The Bleeding Edge

// Episode W31 · 2026-07-24 to 2026-07-31

The frontier kept accelerating this week — and the money and politics underneath it started to crack

The frontier kept accelerating this week — and the money and politics underneath it started to crack. Anthropic shipped Claude Opus 5, its new flagship, and cut over 80% of Claude Code's system prompt to match it; ChatGPT crept up on 1 billion weekly active users. In the same sev…

The Bleeding Edge — Episode Briefing W31

Date range: 2026-07-24 to 2026-07-31 (Europe/Madrid)

Headline of the Week

The frontier kept accelerating this week — and the money and politics underneath it started to crack. Anthropic shipped Claude Opus 5, its new flagship, and cut over 80% of Claude Code's system prompt to match it; ChatGPT crept up on 1 billion weekly active users. In the same seven days, Leopold Aschenbrenner's Situational Awareness — the hedge fund built on his own AGI-is-imminent thesis, which had swollen to as much as $45B — took steep AI losses and had its stock book bought by Citadel at a discount. Sam Altman and Dario Amodei, the two loudest accelerationists in the industry, put their names to a petition asking Washington to build tools to "pace" AI development. The Treasury threatened sanctions over claims that China's Moonshot distilled Anthropic's Fable model, and the US and China scheduled formal AI talks for September. The pattern: the capability curve and the capital-and-geopolitics curve are diverging — the same people who built the "AGI is coming" trade are now either losing money on it or asking the government to slow it down.

Top 5

  1. Anthropic ships Claude Opus 5 — and deletes 80% of Claude Code's system prompt to fit it. Anthropic released Claude Opus 5, its new flagship, positioned at near-frontier intelligence, and disclosed it had removed more than 80% of Claude Code's system prompt because the model no longer needs the scaffolding to behave. Why it matters: the shrinking prompt is the real signal — capability is being absorbed into the weights, so the elaborate instruction-engineering that defined 2024–2025 is becoming a liability rather than a moat. Enterprises that hard-coded workarounds for older models should expect to delete, not add. Corroborated Sources: AI Search, The AI Opportunities on the Claude Code prompt cut.

  2. Aschenbrenner's Situational Awareness fund blows up; Citadel buys the wreckage at a discount. Leopold Aschenbrenner's AI-thesis hedge fund — which had grown as large as ~$45B — took steep losses, and Citadel bought only the fund's stock portfolio, securing a substantial discount. Why it matters: this is the first marquee blow-up of the "long AGI" trade, and it's the author of the Situational Awareness manifesto himself. For executives, it's the clearest evidence yet that being right about the technology and right about the trade are two different things — AI conviction is not a hedge against AI-concentrated drawdowns. Corroborated Sources: WSJ, NYT, CNBC.

  3. ChatGPT crept up on 1 billion weekly active users. Per reporting from The Information, ChatGPT is now approaching 1 billion weekly active users — a scale reached faster than any consumer product in history. Why it matters: at a billion weekly users, ChatGPT is no longer a product category, it's infrastructure — the default interface a meaningful slice of the planet now reaches for first. It also reframes the cost debate: the same week the industry fretted about AI's compute bill, the demand side kept compounding regardless. Corroborated Sources: The Information via The Neuron.

  4. Altman and Amodei back a petition asking Washington to "pace" AI development. Sam Altman and Dario Amodei lent their names to Pacing the Frontier, a petition calling on Washington to build the tooling and institutions needed to deliberately pace frontier AI development. Why it matters: when the two CEOs racing hardest to the frontier publicly ask the government for a brake, read it two ways — a genuine safety signal, and a bid to lock in incumbents by inviting the rules that only the largest labs can afford to comply with. Either way, "self-regulation" just became "please regulate us." Corroborated Sources: Pacing the Frontier, The Neuron.

  5. Treasury threatens sanctions over claims Moonshot distilled Anthropic's Fable — and US–China AI talks are set for September. The US Treasury threatened sanctions over allegations that China's Moonshot AI distilled Anthropic's Fable model to train its own systems, while separately the US and China agreed to hold formal AI talks in September. Why it matters: model distillation is now a trade-policy weapon, not just an engineering technique — the frontier is being treated as strategic export-controlled IP. For any company with a China footprint, model provenance and training-data lineage just became a compliance question, not a research footnote. Unverified Source: Creators' AI weekly digest.

Categorised News

Frontier & Big Tech

Google confirms it's already pretraining Gemini 4 — while Gemini 3.5 Pro is still in testing. Google acknowledged that pretraining on Gemini 4 is underway even though Gemini 3.5 Pro hasn't shipped, and a "Gemini 3.6" reference surfaced separately this week. The takeaway isn't the version number — it's the tempo: Google is running two-plus generations in the pipeline simultaneously, a cadence only a hyperscaler with captive TPUs can sustain. Unverified Sources: Creators' AI digest, AI Search.

Microsoft releases MAI-Cyber-1-Flash, a 5B-active cyber model hitting 95.95% on CyberGym. Microsoft's in-house AI group shipped a security-tuned model with only ~5B active parameters that scores 95.95% on the CyberGym benchmark. A small, cheap, specialised model beating far larger generalists on a security eval is the story enterprises should watch — the economics of defensive AI tooling just improved. Unverified Source: MarkTechPost.

An Anthropic mathematician used Claude to make progress on an open math problem. Per the Creators' AI digest, an Anthropic mathematician used Claude as a working collaborator on an unsolved problem — a concrete data point in the "AI as research partner, not just autocomplete" narrative. Details are thin; treat the framing, not the specifics, as the signal. Unverified Source: Creators' AI digest.

Market Cap / Valuation

Microsoft's Azure cloud revenue rose 43% in the fiscal fourth quarter. Microsoft shares jumped as Azure growth accelerated to 43%, driven heavily by AI workloads. This is the cleanest read yet on where the AI capex is actually landing as revenue — the picks-and-shovels layer is monetising even as the frontier-fund trade (see Aschenbrenner) unwinds. Corroborated Source: Capital Brief.

AI labs set lobbying records — Anthropic spent $1.97M in Q2, beating Nvidia. Anthropic's Q2 federal lobbying spend of $1.97M outpaced Nvidia's, part of a record quarter for AI-lab influence spending in Washington. Paired with the Altman/Amodei pacing petition, it confirms the labs are investing as much in shaping the rules as in shipping models. Unverified Source: Creators' AI digest.

Apps / Dev Tools / Platforms

Perplexity releases pplx, a single-binary CLI that puts search in the terminal. Perplexity shipped pplx, a single-binary command-line tool that brings its search into the terminal — squarely aimed at the agent-and-developer workflow where web access is a primitive, not a browser tab. Unverified Source: MarkTechPost.

Moonshot and kvcache-ai open-source AgentENV, the RL sandbox behind Kimi K3. Moonshot released AgentENV, the reinforcement-learning environment used to train its Kimi K3 agent — notable both as a capable open tool and as a window into how a leading Chinese lab trains agents, published the same week the Treasury threatened it over distillation. Unverified Source: MarkTechPost.

Datalab benchmarks Marker 2 against MinerU, Docling, and LiteParse. Datalab published a head-to-head document-parsing benchmark pitting its Marker 2 against MinerU, Docling, and LiteParse. Document parsing is the unglamorous bottleneck in most enterprise RAG and agent pipelines; a credible benchmark here is more useful to buyers than another chat demo. Unverified Source: MarkTechPost.

Regions / Macro

US GDP grew 1.5% in Q2 as June inflation eased. US second-quarter GDP came in at 1.5% with June inflation softening — a soft-landing-ish backdrop against which the AI-concentrated market stress (Situational Awareness) stands out as idiosyncratic rather than macro-driven. Corroborated Source: Capital Brief.

AI & Robotics

1X pitch deck surfaces: the humanoid backed by OpenAI. A widely-circulated breakdown of 1X's pitch deck framed the humanoid-robotics company's thesis — "for 200 years technology automated cognition; now it automates physical labour" — with OpenAI among its backers. Treat the deck as a fundraising artifact, not a product claim, but the capital flowing into embodied AI is real. Unverified Source: Product Market Fit.

AI Gone Wrong / Disasters / Harms

Researchers broke out of the sandbox in Cursor, Codex, Gemini CLI, and Antigravity. Security researchers demonstrated sandbox escapes across four major AI coding tools — Cursor, Codex, Gemini CLI, and Antigravity — meaning agent code that was supposed to be contained could reach the host. Why it matters: every team running an autonomous coding agent on a developer machine or CI runner is now the trust boundary; "the agent runs in a sandbox" is not a security control you can assume. Unverified Source: Creators' AI digest.

Hugging Face hit by a hack. A security incident at Hugging Face surfaced this week; details in the primary flow are thin. Given how central the platform is to model distribution, teams pulling weights or datasets from it should watch for an official disclosure before assuming supply-chain integrity. Unverified Source: AI Search.

Prompting Skill of the Week

Technique: Subtractive Prompting (Prompt Ablation). Best for: migrating a prompt to a more capable model, or fixing an over-constrained prompt that produces rigid, robotic output. This is the prompt-level version of what Anthropic just did by deleting 80% of Claude Code's system prompt for Opus 5 — newer models need less handholding, and the old scaffolding actively hurts.

  1. Build a test set of 5–10 representative inputs and capture the current output quality with the full prompt as your baseline.
  2. Remove one instruction block at a time — start with the defensive, edge-case, "don't do X" rules, which are the most likely to be obsolete on a stronger model.
  3. Re-run the same inputs. If quality holds, that block was load-bearing on the old model, not this one — leave it out.
  4. Keep cutting until output actually degrades. That's your floor.
  5. Re-add only the specific instructions whose removal broke something concrete.
  6. Write one sentence next to each surviving line explaining why it's still there. If you can't, cut it.

Example prompt (after ablation):

Before: three paragraphs of tone rules, five "never do this" clauses, and a worked example. After: "You are reviewing a customer contract for a mid-market SaaS buyer. Flag anything unusual on liability, auto-renewal, and data rights. Be specific and cite the clause." — and nothing else.

Common failure + fix: you cut an instruction that only matters on rare inputs that aren't in your test set, so the prompt looks fine until a live edge case regresses in production. Fix: keep a separate adversarial eval set of the weird, hostile, and malformed inputs, and ablate against both sets before you ship the leaner prompt.

New AI Tools

pplx (Perplexity). A single-binary command-line tool that puts Perplexity search directly in the terminal. Audience: developers and agent-builders who treat web search as a scriptable primitive rather than a browser destination — drop it into a shell pipeline or an agent's toolset without an SDK. Source: MarkTechPost.

Marker 2 (Datalab). The latest version of Datalab's document-parsing engine, benchmarked this week against MinerU, Docling, and LiteParse. Audience: anyone building RAG or document-agent pipelines who is bottlenecked on turning messy PDFs into clean structured text — parsing quality upstream determines answer quality downstream. Source: MarkTechPost.

AgentENV (Moonshot / kvcache-ai). The open-sourced reinforcement-learning sandbox used to train Moonshot's Kimi K3 agent. Audience: teams and researchers who want to train or evaluate their own agents against a real environment rather than a toy one — and anyone studying how a frontier Chinese lab builds agentic behaviour. Source: MarkTechPost.

AI Personality of the Week

Leopold Aschenbrenner. The former OpenAI researcher who authored the Situational Awareness AGI manifesto and then built a hedge fund of the same name around that thesis had the industry's worst week: after swelling to as much as $45B, the fund took steep AI losses and Citadel bought its stock portfolio at a substantial discount. Aschenbrenner is the purest embodiment of the "AGI is imminent, position accordingly" worldview — and the blow-up is the sharpest available counterexample to it. The technology thesis may still be right; the trade built on it was not durable to the volatility of a market this concentrated. For an executive audience, he's the cautionary case study of the year so far: conviction about where AI is going is not the same as a defensible bet on the timing or the instruments. Sources: WSJ, CNBC.

Catch-All

"Services aren't the new software" — the Sequoia thesis gets a public rebuttal. Sequoia has been telling the market that the next trillion-dollar AI companies will sell work, not software — charge for the outcome (books closed, contracts reviewed, claims processed) instead of a per-seat copilot that competes with every new model release. This week The AI Opportunities published the sharpest counter yet, arguing the "sell the outcome" framing understates how quickly commoditised model capability collapses services margins too. It's the defining strategy debate for anyone deciding whether to build an AI product or an AI-delivered service — and the concrete proof point is stories like GrowthPair, a seven-figure-ARR agency run by three people on Claude. Source: The AI Opportunities.

Show Notes (bullets only)

  • Anthropic ships Claude Opus 5 and deletes 80%+ of Claude Code's system prompt — capability is moving into the weights.
  • ChatGPT crept up on 1 billion weekly active users, per The Information.
  • Leopold Aschenbrenner's ~$45B Situational Awareness fund took steep AI losses; Citadel bought its stock book at a discount.
  • Altman and Amodei backed a petition asking Washington to build tools to "pace" AI development.
  • Treasury threatened sanctions over claims Moonshot distilled Anthropic's Fable; US–China AI talks set for September.
  • Google confirmed it's already pretraining Gemini 4 while Gemini 3.5 Pro is still in testing.
  • Microsoft's MAI-Cyber-1-Flash — a 5B-active cyber model — hit 95.95% on CyberGym.
  • Microsoft Azure cloud revenue rose 43% in the fiscal fourth quarter.
  • Anthropic set an AI-lab lobbying record at $1.97M in Q2, beating Nvidia.
  • Researchers broke out of the sandbox in Cursor, Codex, Gemini CLI, and Antigravity.
  • Perplexity shipped pplx, a single-binary CLI that puts search in the terminal.
  • The Sequoia "services are the new software" thesis drew a sharp public rebuttal.

Weekly Patterns (Inference)

  1. Inference Prompts are shrinking as models grow. Opus 5's 80% Claude Code prompt cut is the leading edge — the instruction-engineering scaffolding built for 2024–2025 models is now dead weight, and migrating to a stronger model increasingly means deleting prompt, not adding it.
  2. Inference The "long AGI" trade is separating from the AGI thesis. Aschenbrenner's fund can blow up in the same week ChatGPT nears a billion users and Opus 5 ships — being right about the technology and right about the trade are now visibly different bets.
  3. Inference The accelerationists are asking for a brake. Altman and Amodei signing a "pace the frontier" petition, plus record lobbying spend, points to a coordinated shift from "self-regulate" to "regulate us" — which conveniently favours the incumbents who can afford compliance.
  4. Inference Model provenance is becoming trade policy. Treasury sanctions threats over Moonshot allegedly distilling Anthropic's Fable turn distillation into an export-control question — expect training-data lineage to become a real compliance surface for any company touching China.
  5. Inference The picks-and-shovels layer is where the revenue is landing. Azure +43% and Microsoft's cheap specialised cyber model contrast with the frontier-fund blow-up — infrastructure and applied tooling are monetising while the speculative frontier trade takes losses.
  6. Inference Agent security is the next enterprise headache. Sandbox escapes across four major coding tools mean "the agent is contained" is now an assumption to verify, not a control to rely on — the developer machine is the trust boundary.
  7. Inference Small specialised models keep beating big generalists on narrow evals. MAI-Cyber-1-Flash at 5B active params scoring 95.95% on CyberGym is another data point that the economics of vertical AI favour tuned, cheap models over frontier general ones for defined tasks.
  8. Inference The product-vs-services strategy question is now the central one. The Sequoia debate, GrowthPair's three-person seven-figure agency, and shrinking prompts all point the same way — as capability commoditises, the durable value is in owning the outcome and the workflow, not the model access.

// Deep dives from this episode