// Episode W23 · 2026-05-30 to 2026-06-06
The IPO and the recursion went public in the same seven days
The IPO and the recursion went public in the same seven days. On 1 June, Anthropic confidentially submitted an S-1 to the SEC at a $965B post-money valuation — the first frontier AI lab to formally enter the public-equity market — and in the same week published evidence that Clau…
The Bleeding Edge — Episode Briefing W23
Date range: 2026-05-30 to 2026-06-06 (Europe/Madrid)
Headline of the Week
The IPO and the recursion went public in the same seven days. On 1 June, Anthropic confidentially submitted an S-1 to the SEC at a $965B post-money valuation — the first frontier AI lab to formally enter the public-equity market — and in the same week published evidence that Claude wrote more than 80% of its own production code in May, including a 52× speedup on the model-training code the lab will rely on to ship its next generation. Microsoft used Build 2026 to declare effective independence from OpenAI with seven proprietary MAI models, GitHub Copilot's switch to token billing punched 10–50× cost spikes through paying customers, Cognition put a $10M "we'll refund Devin if it isn't worth it" productivity guarantee on the table, and SpaceX's $75B IPO closed oversubscribed while Alphabet upsized its equity raise to $85B and Meta openly weighed a mega share sale to fund AI. The pricing of the next-generation labs, the proof those labs are now substantially training themselves, and the macro liquidity to absorb the paper all landed inside the same seven days — and they are pricing the same bet.
Top 5
-
Anthropic files confidential S-1 at $965B valuation. Anthropic submitted a confidential draft S-1 to the SEC on 1 June, targeting an October 2026 listing window. The filing discloses $47B+ ARR, a $2.5B Claude Code run-rate, and that Claude Code is now associated with roughly 4% of public commits on GitHub. Why it matters: the first frontier lab to formally enter the public markets — beating OpenAI to the listing — drags the entire "private super-cap AI lab" category into quarterly disclosure for the first time. From the moment the S-1 is public, frontier-lab gross margins, training-spend amortisation, and customer concentration stop being analyst guesswork. Corroborated Sources: Anthropic Recursive Self-Improvement post (1 Jun), Capital Brief Standup (1 Jun), Creators' AI weekly digest (5 Jun).
-
Claude wrote 80%+ of Anthropic's production code in May; engineers merge 8× more code per day than in 2024. Anthropic published the cleanest internal-productivity disclosure any frontier lab has produced, reporting that more than 80% of production code merged into the Anthropic codebase in May 2026 was authored by Claude, that engineers now merge roughly 8× more code per day than at the 2024 baseline, that internal coding-task success on open-ended problems is at 76% (up 50 points in six months), and that Claude Mythos Preview delivered a ~52× speedup on selected model-training code versus ~3× for Claude Opus 4. Why it matters: the lab pricing itself at $965B in an S-1 published evidence in the same week that its own AI is now doing a large majority of its own engineering — and getting an order-of-magnitude speedup on the recursive workload (training the next model). The recursive-self-improvement argument moves from theory to first measurable corporate disclosure; the same number is now part of both the IPO pitch and the safety debate. Corroborated Source: Anthropic Recursive Self-Improvement (1 Jun).
-
The synchronous IPO + mega-raise rush: SpaceX $75B oversubscribed, Alphabet upsizes to $85B, Meta weighs share sale, OpenAI preparing S-1. Capital Brief reported on 5 June that SpaceX's $75B IPO closed oversubscribed and Meta is weighing a mega share sale specifically to fund AI capex (FT-sourced), on 4 June that Goldman Sachs is pitching a 100-fold AI revenue growth story to clients to support the supply, and on 3 June that Alphabet upsized its historic equity raise to $85B. Bloomberg reported OpenAI is preparing its own confidential S-1 in parallel. Why it matters: more than $2 trillion of implied AI-correlated equity paper is heading toward the public market inside a single calendar quarter, before any of the headline labs file audited public financials. The macro question for H2 2026 is no longer "will the AI labs IPO" — it is whether the public bid absorbs that supply at the private-market multiple, or repriced. Corroborated Sources: Capital Brief Standup (5 Jun), Capital Brief Standup (3 Jun), Bloomberg (1 Jun).
-
GitHub Copilot's switch to token billing torches its own power users — 10–50× cost spikes, $29 → $750/month projections. GitHub retired flat-rate Premium Request Units on 1 June, moving Copilot Pro and Pro+ users to per-token billing for chat, agent mode, and code review. Heavy users are publishing month-over-month bill projections going from $29 to roughly $750. In the same week Microsoft cancelled its own internal Claude Code licenses, Uber capped AI tool spend company-wide, and Cognition launched an enterprise "Devin productivity guarantee" of up to $10M in usage credit if Devin doesn't deliver more engineering value than the customer pays. Why it matters: the pricing model that made developer-AI viral — flat seat fees with hidden inference subsidy — has now visibly broken on the cost side of three different vendors in the same week. The market is bifurcating into "meter everything" (GitHub) versus "outcome-guaranteed" (Cognition), and enterprises are starting to push back hard on the first model and lean toward the second. The Copilot revolt is the first time AI dev-tool pricing has been the leading story of the week. Corroborated Sources: TechCrunch on Copilot token-billing backlash (30 May), Bloomberg on Uber capping AI tool usage (2 Jun), Cognition AI Productivity Guarantee.
-
Microsoft Build 2026: seven proprietary MAI models ship, including MAI-Thinking-1 matching Claude Opus 4.6 on SWE Bench Pro at 10× token efficiency. Microsoft used Build 2026 (2–3 Jun) to launch MAI-Thinking-1 (35B active parameters, 256K context, matches Claude Opus 4.6 on SWE Bench Pro at ~10× lower token cost per task), MAI-Code-1 (coding-specialised), MAI-Image-2.5, MAI-Transcribe-1.5 (43 languages), MAI-Voice-2 (15+ languages), Microsoft Scout (the consumer agent), and surfaced Majorana 2 quantum processor progress with a 2029 scalable-quantum target. Why it matters: Microsoft has now publicly shipped its own competing model in seven categories simultaneously, while internally retiring Claude Code seats. The $13B OpenAI partnership has been load-bearing for Microsoft's AI narrative since 2023; on the evidence of W23 it has stopped being load-bearing for Microsoft's AI products. The clean hyperscaler-owns-the-stack story moves from rumour to product launch. Corroborated Source: Microsoft Build 2026 keynote post.
Categorised News
Market Cap / Valuation
Suno raises $400M at $5.4B valuation despite active copyright litigation. AI music generator Suno raised $400M led by Bond Capital at a $5.4B post-money on 3 June, doubling its valuation in seven months while Universal, Sony, and Warner copyright lawsuits remain active. ARR is reported at $300M and the platform generates roughly 7M songs per day. Why it matters: the music-rights litigation has not slowed the funding round and the disclosed ARR is the cleanest single read on consumer-AI willingness-to-pay for generative-audio. Corroborated Source: TechCrunch (3 Jun).
Supabase raises $500M at $10.5B valuation. Open-source Postgres-backed database platform Supabase closed a $500M round at a $10.5B valuation on 4 June, with the round explicitly positioned as funding the AI-agent-backend expansion. Why it matters: the agent-backend layer (auth, storage, RLS, vector, edge) is consolidating into Postgres-superset platforms; Supabase's print is the second $10B+ round in that category since W19. Corroborated Source: CNBC (4 Jun).
SoftBank commits up to €75B for French AI data-centre buildout. SoftBank announced up to €75B (≈$87.5B) for AI data centres in France on 1 June, €45B in Phase 1, 5 GW total capacity with 3.1 GW concentrated in Hauts-de-France, 2031 completion target. Why it matters: Europe's largest single AI infrastructure commitment to date and the cleanest hyperscaler-equivalent signal that French nuclear baseload + regulatory frame is now considered competitive with US dollar-priced power. Corroborated Source: CNBC (31 May).
TSMC: AI chip demand will outstrip supply "for a long time." TSMC CEO C.C. Wei told shareholders this week that meeting current AI-driven demand will take "a long time" while pledging to hold pricing stable rather than implement headline price hikes. Why it matters: the supply-side constraint on frontier training and serving is being publicly extended by the only company that can credibly characterise it; combined with SK Hynix crossing $1T market cap in W22, the AI infrastructure rent is still accruing predominantly to East Asian manufacturing capacity. Corroborated Source: Tom's Hardware.
Frontier & Big Tech
ChatGPT crosses 1B monthly active app users; Claude hits 56M MAUs at 640% YoY growth. Sensor Tower data indicates ChatGPT became the fastest app in history to 1B monthly active users in May 2026, growing 62% YoY, while Claude reached 56M MAUs at 640% YoY growth. The same dataset shows US Claude installers reduce subsequent-month ChatGPT time by ~5%. Why it matters: the consumer-AI market is segmenting, not consolidating — Claude's growth rate at the smaller base implies meaningful share-of-attention erosion at the upper end, not just net new addressable users. The winner-take-all framing is now empirically wrong at the consumer layer. Unverified Source: TheNextWeb.
Meta Business Agent ships globally on 3 June across WhatsApp, Instagram, and Messenger. Meta released Business Agent to over 1M businesses globally on 3 June, enabling no-code deployment of autonomous AI agents for sales, support, and operations across WhatsApp, Instagram, and Messenger, with first-party integration to Shopify and Zendesk. Meta cites 1B+ daily business-to-customer messages already flowing through the platform layer. Why it matters: Meta has converted the messaging-layer install base into a distribution channel for agents in one move, with a free-to-start pricing model that bypasses both the OpenAI enterprise stack and the Salesforce/HubSpot CRM stack. The CRM-vs-Messaging frontline of agent distribution just shifted. Corroborated Source: Meta newsroom (3 Jun).
OpenAI turns Codex into an enterprise automation platform. OpenAI converted Codex on 2 June from a coding tool into an enterprise automation surface with web app deployment, 62 business app plugins, and 110 pre-built skills — explicitly competing with Zapier, Make, and ServiceNow. Codex is reported at 5M weekly users with 20% non-developers and non-developer adoption growing 3× faster than developer adoption. Why it matters: the agent-platform consolidation thesis (one prompt, every SaaS connector) just moved from "coming soon" to shipped at OpenAI scale, and the non-developer growth curve is the structural reason Zapier and ServiceNow have to defend on price. Corroborated Source: TheNextWeb.
OpenAI ships ChatGPT memory with reviewable summaries. OpenAI deployed an upgraded ChatGPT memory system this week that builds longitudinal user context and exposes a reviewable summary the user can edit or wipe. Why it matters: the EU AI Act, the Italian DPA and W22's UK opt-out ruling are all moving in the direction of explicit user-facing data controls; OpenAI shipping a reviewable memory layer is the cleanest signal that frontier-lab product policy is now front-running regulation rather than reacting to it. Corroborated Source: OpenAI (ChatGPT Memory).
Claude Opus 4.8 ("Mythos") finishes its global rollout — including Australia. Capital Brief confirmed on 2 June that Australian customers now have access to Claude Opus 4.8, completing the global rollout of the model first announced in W22. Corroborated Source: Capital Brief Standup (2 Jun).
Apps / Dev Tools / Platforms
Cognition launches AI Productivity Guarantee for Devin — up to $10M in coverage. Cognition committed in writing to fund Devin usage up to $10M per enterprise customer if Devin doesn't deliver engineering value greater than the contracted spend, calculated as auditable estimated time savings versus human labour. Why it matters: the first major dev-AI vendor to put a contractual outcome floor on the table, directly into the teeth of GitHub Copilot's token-billing pricing event in the same week. The pricing arc is bending from "rent the seat" to "guarantee the outcome" in 2026. Corroborated Sources: Cognition blog, Devin Guarantee page.
Microsoft cancels internal Claude Code licences; Uber caps AI tool usage. Bloomberg reported on 2 June that Microsoft has stopped paying for internal Claude Code seats and that Uber engineering leadership issued a memo capping AI tool usage company-wide. Why it matters: large-customer pull-back is the leading indicator the dev-AI pricing model needs to reprice; this is the first time a hyperscaler and a marquee enterprise have done it publicly in the same week. Corroborated Source: Bloomberg (2 Jun).
Infrastructure & Ecosystem
NVIDIA releases Nemotron 3 Ultra: 550B parameter open-weight MoE Mamba-Transformer with 1M context. Nemotron 3 Ultra is a hybrid Mamba-Transformer mixture-of-experts model targeted at long-running agentic workloads. The technical report claims extended-context inference economics that materially undercut equivalent transformer-only stacks at the same effective context. Why it matters: the long-running-agent inference economics story has been NVIDIA's missing piece versus the closed labs; 1M context at 550B active-weights-equivalent open under MIT-class licence collapses the moat on a class of agent workloads that dominated the W22 + W23 enterprise narrative. Corroborated Sources: MarkTechPost (4 Jun), NVIDIA technical report PDF.
NVIDIA Cosmos 3: a two-tower mixture-of-transformers foundation model unifying physical reasoning, world generation, and action generation. Cosmos 3 lands as the foundation model for the robot-and-world-sim track — the architecture explicitly separates a physical-reasoning tower from an action-generation tower under a shared mixture-of-experts router. Why it matters: the world-model layer the humanoid-robot category has been waiting on now exists from NVIDIA; combined with the 1X World Model Lab launch the same week, the embodied-AI stack is moving from research demos to released infrastructure. Corroborated Source: MarkTechPost (3 Jun).
Google ships Gemma 4 12B for laptop-local inference + Gemma 4 QAT mobile checkpoints. Google AI Edge released Gemma 4 12B for direct laptop execution on 5 June (no cloud round-trip required) and shipped Q4_0 QAT checkpoints with a new mobile-optimised format. Why it matters: the on-device frontier is moving on a roughly monthly cadence now (Gemma 4 12B laptops, Q4_0 mobiles, Apple-silicon-native MisoTTS) — and the privacy-sensitive enterprise segment now has credible local options that match cloud quality on a meaningful slice of tasks. Corroborated Sources: Google Developers Blog (5 Jun), MarkTechPost on Gemma 4 QAT.
Miso Labs releases MisoTTS — open-weight 8B emotive text-to-speech. MisoTTS lands with open weights, 8B parameters, emotive prosody control. Pairs naturally with the OSCAR + EAGLE 3.1 W22 serving stack improvements for low-latency voice agents. Corroborated Source: MarkTechPost (4 Jun).
OpenJarvis (Stanford SAIL + Hazy Research): local-first framework for on-device personal AI agents — re-surfaced this week. OpenJarvis decomposes a personal AI into five swappable primitives (Intelligence, Engine, Agents, Tools & Memory, Learning) serialised into a portable TOML "spec," runs fully on-device via Ollama/vLLM/llama.cpp, and wires up 25+ data connectors and 32+ messaging channels. Originally released by Stanford in March 2026 (v0.1.0, Apache 2.0); re-featured by MarkTechPost on 3 June, which is how it landed in this week's pull. Why it matters: combined with Gemma 4 12B local + Nemotron 3 Ultra open, the "agent runs entirely on your laptop" stack is now real — but see the Deep Dive below for the gap between the architecture and actual production adoption. New this week: no — re-feature of a March release, not a fresh launch. Corroborated Sources: MarkTechPost (3 Jun), Stanford Scaling Intelligence Lab.
TinyFish BigSet: open-source multi-agent system that builds structured live datasets from plain-English descriptions. BigSet replaces the scraper-and-pipeline pattern with an agent crew that maintains a live, structured dataset on demand. Useful pattern primitive for the agentic-data-engineering stack. Corroborated Source: MarkTechPost (2 Jun).
StepFun ships Step 3.7 Flash multimodal model; OpenScientist releases AutoScientists; NVIDIA LocateAnything paper. Three faster-moving research drops surfaced in the AI Search weekly roundup: StepFun's Step 3.7 Flash (efficiency-tilted multimodal for agent frameworks with reliable tool calls), OpenScientist's AutoScientists (multi-agent self-organising research-team framework), and NVIDIA's LocateAnything (natural-language object tracking in video for autonomous robotics and search). Unverified Sources: AI Search weekly (31 May), NVIDIA LocateAnything, StepFun blog, AutoScientists.
Regions / Macro
Trump signs AI Innovation and Security Executive Order. President Trump signed an executive order on 2 June directing federal agencies to accelerate AI adoption, expand training-dataset access, streamline procurement, and designate the Department of Commerce as the lead AI standards coordination agency. Why it matters: the US federal posture explicitly diverges from the EU AI Act compliance frame, with Commerce — not OSTP, not a new AI agency — owning the standards lane. The procurement-acceleration clauses are the first material federal demand-side commitment to enterprise AI in 2026. Corroborated Source: White House Presidential Actions (2 Jun).
UK regulators mandate Google AI Search opt-out for publishers. UK regulators ruled that Google must give publishers an opt-out mechanism for generative AI Search features. Rollout begins in the UK and is expected to expand globally. Why it matters: the publisher-versus-AI-overview dispute now has a regulated opt-out path in one major market — the precedent will be cited heavily by EU news publishers and Australian AP-style content groups within 2026. Corroborated Source: TechCrunch (3 Jun).
AI & Robotics
1X launches Humanoid Robot World Model Lab. Robotics firm 1X announced a dedicated World Model Lab on 4 June, with founder Bernt Børnich arguing that general-purpose humanoid robots require purpose-built world models rather than fine-tuned vision-language models. Why it matters: combined with NVIDIA Cosmos 3 shipping the same week, the world-model layer is now the explicit research frontline for embodied AI. The fine-tune-only humanoid approach (the dominant 2024–2025 method) is being publicly retired by category leaders. Corroborated Source: Forbes (4 Jun).
AI in Consumer Hardware
NVIDIA returns to the PC market with a new AI chip + on-device agent positioning. Capital Brief reported on 1 June that NVIDIA is re-entering the PC chip market with an AI-targeted SoC, and the Neuron Daily flagged a coordinated NVIDIA + Microsoft push to position next-generation AI coworkers as primarily on-device rather than cloud-hosted. Why it matters: the on-device-agent thesis that drove the Gemma 4 12B + OpenJarvis releases this week now has a coordinated silicon + OS commitment behind it. The competitive line between consumer AI hardware (Apple Neural Engine, Qualcomm Snapdragon X, NVIDIA-PC) is sharpening. Unverified Sources: Capital Brief Standup (1 Jun), The Neuron Daily (2 Jun).
AI Gone Wrong / Disasters / Harms
GitHub Copilot 10–50× cost spike: "What a joke." The TechCrunch headline ("What a joke: GitHub Copilot's new token-based billing spurs consternation among devs") and the developer-Twitter reaction are the cleanest single-week customer revolt against AI pricing on record. Corroborated Source: TechCrunch (30 May).
Anthropic's recursive-self-improvement claim splits the safety community. The 80%-of-code disclosure and the 52× speedup number prompted immediate pushback from interpretability researchers arguing the metrics conflate "Claude wrote the diff" with "Claude designed the change," and from safety researchers concerned the framing legitimises capability acceleration arguments. Anthropic published the post the same day as the S-1 filing. Inference Source: editorial read of the Anthropic post and W23 developer-Twitter response — no single corroborating outlet yet.
Prompt-injection / "your AI assistant can be hijacked right now." The Neuron Daily flagged a broader prompt-injection security advisory cycle this week tied to the proliferation of agent deployments without input/output sanitisation. Unverified Source: The Neuron Daily (4 Jun).
Prompting Skill of the Week
Technique: The Productivity Guarantee Pattern. Best for: deciding whether to deploy an AI agent for a specific job-to-be-done in your team or company. Inspired by Cognition's $10M Devin productivity guarantee — adapted from a vendor contract into a personal/team evaluation prompt.
The pattern forces the model to do the work Cognition's contract does: define the productivity calculation upfront, with auditable inputs, before the work begins.
- Describe the recurring task you want to delegate to an AI agent, in plain language. Include the current human-time cost and a sample of the inputs.
- Ask the model to enumerate, before doing the work, the explicit assumptions it will make to claim productivity gain — e.g. "I assume the input format is consistent week-to-week," "I assume the human baseline is 45 minutes."
- Ask the model to write the verification it would need to confirm the productivity claim — e.g. "Compare the agent's draft against the human's draft for X criteria; rate match-rate."
- Run the agent on three real samples. Apply the verification it pre-wrote.
- If verification fails, the model rewrites step 2 (its assumptions) and you re-run. If verification passes on all three, the agent is provisionally deployed with the verification as the standing audit.
Example prompt (step 2):
"Before producing any output, enumerate every assumption you will make about my inputs, my baseline human process, and the quality standard I require. Number each assumption. For each, state what I should check to confirm it holds. Do not start the work until I confirm the assumption list."
Common failure + fix: the model produces a beautiful output that passes a vibes-check but doesn't survive the audit step — because the audit was never specified. Fix: do not skip step 3. The audit is the productivity guarantee; without it you are buying a seat licence, not a result.
Variant — team rollout: in step 4, run the same prompt across three different team members' samples. If the agent's match-rate varies more than 15 points across team members, the agent is not yet deployable as a shared tool — only as an individual workflow.
New AI Tools
Gemma 4 12B (Google AI Edge, laptop-local). 12B-parameter model running fully on-device on consumer laptops — no cloud round-trip, no data egress — with QAT-quantised Q4_0 mobile checkpoints alongside it for handsets. Audience: privacy-sensitive enterprise teams, on-prem regulated industries, and any individual who wants their agent's context to never leave the device. Quick workflow: install Google AI Edge runtime, point Gemma 4 12B at a local document folder, build a private RAG-over-laptop in under thirty minutes. Source: Google Developers Blog (5 Jun).
Nemotron 3 Ultra (NVIDIA, 550B MoE Mamba-Transformer, 1M context, open). Long-running agent workloads that previously required Claude or GPT-class closed APIs now have a credible open option. Audience: any team building agentic workflows that need 100K+ context inference economics. Quick workflow: spin up on NVIDIA inference stack, drop into existing agent harness as a closed-model substitute, benchmark against your current cost-per-task. Sources: MarkTechPost (4 Jun), Technical report PDF.
OpenJarvis (Stanford SAIL + Hazy Research, local-first personal-agent framework). Apache-2.0 framework for personal-agent stacks that run offline by default, cloud-optional — tools, memory, and an on-device learning loop. Audience: developers who want a personal-agent skeleton that doesn't assume an OpenAI or Anthropic API key. Quick workflow: git clone, uv sync, jarvis init (auto-detects hardware + recommends a model), jarvis doctor, then run the morning-digest-mac preset against your own Gmail + Calendar. See the full Deep Dive below. Source: MarkTechPost (3 Jun).
AI Personality of the Week
Mustafa Suleyman. Suleyman runs Microsoft AI and is the named force behind the Build 2026 MAI launch — seven proprietary Microsoft models shipping simultaneously, including MAI-Thinking-1 matching Claude Opus 4.6 on SWE Bench Pro at roughly an order of magnitude less per-task cost. He co-founded DeepMind in 2010, was bought by Google in 2014, founded Inflection AI in 2022, and joined Microsoft in March 2024 when Microsoft hired him and most of the Inflection team into a newly stood-up Microsoft AI division. The W23 move is the culmination of that hire: Microsoft has gone from "OpenAI's largest customer" to "competitor with OpenAI in seven categories simultaneously" inside fourteen months. The same week, the company quietly stopped paying for internal Claude Code licences. The signal is unambiguous — Microsoft has decided the OpenAI partnership is no longer load-bearing for its AI product strategy, and Suleyman is the executive who got it there. The second-order question to watch: where Suleyman's research-recruit pull lands now that Karpathy is at Anthropic and Microsoft is publicly competing in his old reporting line at Google DeepMind. Source: Microsoft Build 2026.
Catch-All
Benedict Evans on Lenny's Podcast: "AI is a 1997 internet moment." Benedict Evans argued in a long-form interview that 2026 maps to 1997 in internet-history terms — meaning real revenue, real adoption, real public-market participation, and a real distance still to run before the structure of the eventual winners is visible. The contrarian take inside the W23 IPO rush is exactly his: if you believe the 1997 analogy, the Anthropic + OpenAI listings are pricing the equivalent of Yahoo + Excite pre-dotcom-bust, with the post-bust Google equivalent not yet visible in the W23 capital tables. Why it matters: the most-cited bearish frame for the AI capex cycle now has a clean public articulation from a credible analyst inside the largest distribution podcast for tech operators. Expect it to be repeated in board decks and earnings calls through the IPO window. Unverified Source: Lenny's Podcast — Benedict Evans interview (31 May).
Deep Dive: OpenJarvis — "Personal AI, on Personal Devices"
Cold open. The same week Anthropic filed to go public at $965B on the strength of Claude writing its own code inside a data centre, a Stanford lab's answer to the question "where should your personal AI actually live" quietly crossed 5,000 GitHub stars — running on a Mac Mini, calling the cloud zero times. OpenJarvis is the counter-narrative to this week's headline: not a bigger central brain you rent, but a portable one that fits on your desk.
What it actually is
OpenJarvis is an open-source (Apache 2.0) framework from Stanford's Scaling Intelligence Lab and Hazy Research — the groups associated with Christopher Ré and Azalia Mirhoseini — for building personal AI agents that run on your own hardware, cloud-optional rather than cloud-default. It was released in March 2026 (v0.1.0, arXiv 2605.17172) and re-surfaced in this week's newsletters via a MarkTechPost re-feature, so despite landing in the W23 pull it is not new this week. By June it sits at roughly 5.4k stars / 1.2k forks (Python ~83%, Rust ~9%, TypeScript ~7%).
The thesis the authors lead with is a historical analogy: personal AI in 2026 is where computing was in the late-1970s mainframe-to-PC shift. Their "Intelligence Per Watt" study is the evidence — local models can accurately handle 88.7% of single-turn chat and reasoning queries at interactive latency, and local intelligence efficiency improved 5.3× between 2023 and 2025. The claim isn't "local is more powerful." It's "local is now good enough for the overwhelming majority of what a personal agent actually does — so why are you renting it from someone else's server?"
The architecture in one breath
OpenJarvis breaks a personal AI into five swappable, typed primitives, serialised into a single portable TOML file called a spec:
| Primitive | What it controls | Examples |
|---|---|---|
| Intelligence | model, weights, quantisation, gen params | Qwen3.5, Gemma 4, Nemotron, Granite |
| Engine | inference runtime | Ollama, vLLM, SGLang, llama.cpp, Apple Foundation Models, Exo (+5 cloud backends) |
| Agents | the reasoning loop | ReAct, CodeAct, tool-use policy |
| Tools & Memory | the outside world | 25+ data connectors, 32+ messaging channels, native MCP, swappable memory |
| Learning | self-improvement | LoRA, DSPy, GEPA, LLM-guided spec search |
The interesting primitive is Learning. OpenJarvis's "LLM-guided spec search" uses a frontier cloud model once, at optimisation time, as a teacher: it reads your agent's failure traces, proposes edits across all five primitives, and keeps only the edits that fix the targeted failures without regressing anything else (a "gate," 1% default tolerance). The optimised spec then runs fully on-device with zero cloud calls at inference. You pay for the cloud teacher during tuning; you never pay for it again at runtime.
How people are actually using it (the honest version)
Here's the candid part the marketing won't lead with: as of early June there are essentially no documented production users or testimonials. Both the creators' page and the independent reviews describe intended use, not deployed use — this was a six-days-old research release at first review. So "how people use it" really means "what it ships ready to do." Out of the box it includes eight built-in agents across three execution modes (on-demand, scheduled, continuous), with presets you can run in about three minutes:
morning-digest-mac— a spoken daily briefing (text-to-speech) assembled from your Gmail + Calendardeep-research— multi-hop research with citationscode-assistant— an agent with code execution and shell accessscheduled-monitor— a stateful agent that wakes on a cron schedule
The connector list is the practical hook: Gmail, Calendar, iMessage, Notion, Obsidian, Slack, GitHub on the input side; WhatsApp, Telegram, Discord, iMessage, Signal as the chat surfaces you talk to it through. It can import ~150 Hermes Agent skills and ~13,700 OpenClaw community skills, so the skill library isn't starting from zero even while the OpenJarvis-native ecosystem is thin.
The numbers that matter
On the authors' 8-benchmark / 508-task suite, the best local model (Qwen3.5-122B) hits 80.3% average accuracy versus Claude Opus 4.6 at 83.5% — a 3.2-point gap — and matches or beats cloud on 4 of 8 benchmarks (tool-calling, coding, agentic τ-Bench). The portability result is the strongest selling point: naively swapping a cloud model for a small local one (Qwen3.5-9B) in competing frameworks tanks accuracy 25–39 points; under an OpenJarvis spec the drop shrinks to 5.6–16.5 points — recovering 56–77% of the portability loss. Cost/latency claims: roughly 800× lower marginal cost per call and ~4× lower latency than the cloud equivalent, with energy telemetry sampled every 50ms.
Honest take — what works and what doesn't
- Works: the local-first design is real, not a checkbox ("offline by default, cloud by choice"). The five-primitive decomposition is the most thoughtful local-agent architecture published to date. Energy and cost as first-class metrics is genuinely novel. The on-device learning loop is unique — no competing framework offers it.
- Doesn't (yet): it's v0.1.0 with rough edges (macOS needs an
xattr -crunstick on first launch). The learning loop has no published benchmarks — nobody has shown the on-device fine-tuning actually improves real-world task quality. The ecosystem is bare next to LangChain's integrations and production war stories. There's a real hardware floor: 8GB VRAM minimum, 16GB+ for the learning loop — "most laptop users hit walls quickly." And the 11.3% of queries local can't handle are exactly the hard reasoning/research tasks, so you still keep a cloud fallback.
Actionable takeaway
To pilot it: git clone, uv sync, jarvis init (auto-detects hardware, recommends a model), jarvis doctor, then run the morning-digest-mac preset against your own Gmail + Calendar for a week. Go/no-go metric: does the local digest match what you'd have written yourself ≥80% of the time, with zero cloud calls? If yes on a Mac Mini M4 or a 16GB+ GPU box, the personal-agent economics flip in your favour. If you're on a sub-8GB laptop or you need best-in-class reasoning, wait a release.
Alternatives worth naming
Hermes Agent (Nous Research) and OpenClaw are the closest comparables on the personal-agent side — OpenJarvis explicitly imports both their skill libraries and beats both on the model-swap portability test. LangChain / LlamaIndex remain the mature-but-cloud-assuming default. Ollama + a hand-rolled script is the "just do it yourself" baseline OpenJarvis is trying to replace with structure.
Talking points / questions for the show
- Is "your personal AI runs on someone else's server" a problem normal people feel — or only one privacy nerds feel? (This week's Copilot 10–50× bill shock suggests cost, not privacy, may be the wedge that sells local.)
- The recursion mirror: Anthropic's data-centre AI improving its own code (this week's headline) vs OpenJarvis's laptop AI improving its own spec. Same loop, opposite locations. Which one matters more for the average person?
- Stanford built the framework but won't run your agent. Who builds the ecosystem that makes a local framework trustworthy for production — and does "v0.1.0 from a research lab" ever cross that chasm?
Pitfalls to flag on-air
Don't oversell adoption — there's none documented yet; this is a strong artifact, not a movement. Don't conflate "88.7% of queries" with "88.7% as good" — it's a coverage stat, not a quality stat. And the 800× figure is marginal cost — it ignores the hardware you bought and the watts the telemetry is busy counting.
Project tie-in + derivatives
Directly relevant to the show's own "Hermes local LLM" AI-Projects case study (building-with-ai) — OpenJarvis imports Hermes Agent skills and is the cleanest framework for a local-first build. Derivative angles: (1) a "build your own Jarvis on a Mac Mini" YouTube short using the morning-digest-mac preset; (2) a LinkedIn post on the dev-AI pricing recession → local-first as the cost hedge (tie Copilot's 10–50× to OpenJarvis's 800×); (3) a content/projects/ entry documenting a real OpenJarvis + Hermes deployment as proof-of-work for the KPI page.
Shareable Lines
- "Anthropic filed an S-1 at $965B the same week it disclosed Claude wrote 80% of its own production code. The IPO and the recursion are pricing the same bet."
- "GitHub Copilot just torched its own power users with token billing. Cognition responded with a $10M productivity guarantee. The flat-rate-seat era of dev-AI is over."
- "Microsoft went from OpenAI's biggest customer to competing in seven categories in one keynote. The $13B partnership stopped being load-bearing this week."
- "Five companies open-sourced models this week with weight counts adding up to more than 1.6T parameters. Two of them ran on a laptop."
Show Notes (bullets only)
- Anthropic filed confidential S-1 at $965B post-money valuation; targets October 2026 listing; first frontier AI lab to formally enter public markets.
- Anthropic publishes "Recursive Self-Improvement" post the same day — Claude wrote 80%+ of production code in May, engineers merge 8× more code per day than 2024, 52× speedup on model-training code via Claude Mythos Preview.
- SpaceX $75B IPO oversubscribed; Alphabet upsizes equity raise to $85B; Meta weighs mega share sale to fund AI capex; OpenAI preparing parallel confidential S-1.
- GitHub Copilot switches to token billing — power users projecting 10–50× cost spikes ($29 → ~$750/month).
- Cognition launches Devin Productivity Guarantee — up to $10M per customer if Devin doesn't deliver more engineering value than contracted spend.
- Microsoft cancels internal Claude Code licences; Uber caps AI tool usage company-wide.
- Microsoft Build 2026 ships seven proprietary MAI models — MAI-Thinking-1, MAI-Code-1, MAI-Image-2.5, MAI-Transcribe-1.5, MAI-Voice-2, Microsoft Scout, Majorana 2.
- MAI-Thinking-1 matches Claude Opus 4.6 on SWE Bench Pro at ~10× token efficiency, 35B active parameters, 256K context.
- ChatGPT crosses 1B monthly active app users; Claude reaches 56M MAUs at 640% YoY growth — segmentation, not winner-take-all.
- Meta Business Agent ships globally on 3 June across WhatsApp, Instagram, Messenger; 1M+ businesses on early version.
- OpenAI Codex becomes enterprise automation platform — 62 business app plugins, 110 pre-built skills, competing with Zapier/Make/ServiceNow.
- OpenAI ships ChatGPT reviewable-memory upgrade.
- Suno $400M at $5.4B; Supabase $500M at $10.5B; SoftBank up to €75B for French AI data centres.
- NVIDIA Nemotron 3 Ultra: 550B open MoE Mamba-Transformer, 1M context.
- NVIDIA Cosmos 3: two-tower mixture-of-transformers foundation model for physical reasoning + world generation + action generation.
- Google Gemma 4 12B for laptop-local inference; Gemma 4 QAT mobile checkpoints.
- Miso Labs MisoTTS 8B emotive open-weight TTS.
- OpenJarvis: local-first personal-agent framework.
- TinyFish BigSet: multi-agent system building structured live datasets from plain-English descriptions.
- StepFun Step 3.7 Flash; OpenScientist AutoScientists; NVIDIA LocateAnything.
- Trump signs AI Innovation and Security Executive Order — Commerce as lead standards agency.
- UK regulators mandate Google AI Search opt-out for publishers.
- TSMC: AI chip demand will outstrip supply "for a long time"; prices held stable.
- 1X launches Humanoid Robot World Model Lab.
- NVIDIA re-enters PC market with AI-targeted SoC.
- Claude Opus 4.8 ("Mythos") rollout reaches Australia.
- Benedict Evans on Lenny's Podcast: AI = 1997 internet moment; bear case for IPO multiples gets its W23 articulation.
Inference Bullets (weekly patterns + tensions)
-
Inference The Anthropic IPO + recursive-self-improvement post are a single coordinated disclosure event. The valuation depends on the productivity claim being true; the productivity claim depends on the IPO providing the capital to ship the next model. The two documents read together are tighter than either reads apart.
-
Inference Dev-AI is entering its first real pricing recession. Token billing on Copilot, Microsoft killing its own Claude Code spend, Uber capping AI tools, and Cognition's outcome guarantee are all the same trade — the seat-licence model is being repriced toward results pricing inside a single quarter.
-
Inference Microsoft has decided the OpenAI partnership is a cost centre, not a moat. Seven MAI models + internal Claude Code cancellation + the implicit competition with Mustafa Suleyman's old DeepMind tree are not the actions of a partner.
-
Inference The on-device / open-weight cadence is now the dominant infrastructure story of every week. Gemma 4 12B, Nemotron 3 Ultra, MisoTTS, OpenJarvis, Cosmos 3, BigSet, Step 3.7 Flash, AutoScientists, LocateAnything — nine open releases in W23 alone. The "frontier-closed-labs-win-by-default" thesis is being eroded weekly at the infrastructure layer even as the IPO rush prices the opposite story at the brand layer.
-
Inference The IPO rush is exposing a forecast mismatch. Goldman Sachs's pitch of "100× AI revenue growth" is the bull frame; Benedict Evans's "1997 internet moment" interview is the bear frame; the truth is almost certainly that the rush prices in a winner whose name is not yet on the W23 capital table.
-
Inference The hyperscaler-owns-the-stack and outcome-guaranteed pricing trades are the same trade — both are responses to the cost-of-inference reality that the W23 numbers (Copilot 10–50×, MAI-Thinking-1 10× efficiency, Anthropic 52× speedup) finally made visible to non-specialist buyers.
-
Inference The world-model layer is the new research frontline. NVIDIA Cosmos 3 + 1X World Model Lab + AutoScientists in the same week is the cleanest signal that the post-LLM frontier is being defined as "models that can reason about physical and scientific causality," not bigger transformers.
-
Inference The regulatory posture has split clean. Trump's AI Innovation EO (acceleration + procurement) and the UK Google AI Search opt-out (publisher protection) are now the two clearest poles, with the EU AI Act compliance frame sitting between them. Expect this trichotomy to drive enterprise legal-architecture decisions through year-end.
-
Inference Anthropic's quarter-of-the-year (Karpathy join W22, S-1 + Mythos rollout + recursive post W23) is now the cleanest concentrated burst of frontier-lab momentum since GPT-4. The strategic question for OpenAI is whether the parallel S-1 closes the narrative gap or codifies the second-place framing.
-
Inference The W23 capital cycle, the W23 productivity claim, and the W23 model-release cadence are three distinct disclosures by three distinct categories of actor — but they are all reading off the same underlying trade. The signal worth watching across the next four weeks is which of the three breaks first.
Editorial Notes
- Briefing produced manually 2026-06-06 (Saturday); friday-pipeline failed Friday 2026-06-05 at briefing.generator due to
claude -preturning a 2048-char degenerate response from an 88,868-char prompt (ratio 0.02, no YAML frontmatter). Smoke-test on Saturday confirmedclaude -phealthy — the W23 failure was transient. SMTP_APP_PASSWORD secret is also stale (last set 2026-05-09), so the failure was silent to both hosts. - Verification status: all Top 5 stories have primary-source citation. The Anthropic recursive-self-improvement claim is corroborated against Anthropic's own post; downstream interpretive splits are tagged Inference. Two consumer-AI MAU numbers (ChatGPT 1B / Claude 56M) sourced to a single outlet and tagged Unverified pending second-source pickup.
- Articles + Substack newsletter generation not produced in this manual run — those depend on the pipeline's article-writer + newsletter modules. The briefing alone is the recovery artefact; long-form derivatives can be triggered separately once the pipeline is green.