The Bleeding Edge

// Episode W32 · 2026-07-31 to 2026-08-07

The software layer went free this week; the physical layer went national

The software layer went free this week; the physical layer went national. Moonshot dropped full Kimi K3 weights, DeepSeek put V4-Flash-0731 on Hugging Face under an MIT licence, NVIDIA open-released a 34B driving model under OpenMDW-1.1, and Microsoft published research showing t…

The Bleeding Edge — Episode Briefing W32

Date range: 2026-07-31 to 2026-08-07 (Europe/Madrid)

Headline of the Week

The software layer went free this week; the physical layer went national. Moonshot dropped full Kimi K3 weights, DeepSeek put V4-Flash-0731 on Hugging Face under an MIT licence, NVIDIA open-released a 34B driving model under OpenMDW-1.1, and Microsoft published research showing that optimised agent skills transfer across model scales and between the Codex and Claude Code harnesses. In the same seven days, Washington moved to ban foreign-made humanoid robots, Brussels opened a €30B call for seven AI gigafactories at 100K+ chips each, and Seoul signed a reported $950B in AI deals in San Francisco. The thing that used to be the moat — weights, prompts, harness lock-in — is being given away. The thing nobody thought to fight over — bodies, fabs, grid connections, and who is allowed to sell you a robot — is now industrial policy.

Top 5

  1. OpenAI claims 1 billion users and says 99.8% of its tokens are now agentic. The figure circulated via the Creators' AI weekly digest: a billion-user milestone paired with the claim that all but 0.2% of tokens served are consumed inside agentic loops rather than human chat turns. Why it matters: if directionally true, "ChatGPT usage" is no longer the metric that matters — the unit of consumption is an agent run, and every budget, rate-limit, audit trail, and vendor contract written for chat is now mis-specified. Unverified Source: Creators' AI weekly digest.

  2. The open-weight flood: Kimi K3 full weights, DeepSeek-V4-Flash-0731 under MIT, MiniMax H3, and AMD model releases — all inside one week. Moonshot AI published the complete weights for Kimi K3, described in the digest as the first openly downloadable model of its class. DeepSeek shipped V4-Flash-0731 to Hugging Face under an MIT licence — the most permissive terms available, with no use restrictions and no field-of-use carve-outs. MiniMax H3 and new AMD models landed in the same window. Why it matters: MIT-licensed frontier-adjacent weights mean a European or US enterprise can self-host a capable model with no vendor relationship at all — and the labs supplying that option are overwhelmingly Chinese. Corroborated Sources: AI Search weekly, Creators' AI weekly digest.

  3. The US moves to ban foreign-made humanoid robots, aimed squarely at China. Reported this week as a prohibition on foreign-manufactured humanoid robots, framed explicitly as a China measure. Mechanism and scope — outright import ban versus federal procurement exclusion — are not specified in the source. Why it matters: chip export controls took four years to arrive; the humanoid equivalent has arrived before the category has meaningful commercial deployment, which tells you the physical layer is now being fought over pre-emptively rather than reactively. Unverified Source: Creators' AI weekly digest.

  4. Microsoft's SkillOpt shows optimised agent skills transfer across model scales — and between Codex and Claude Code. Microsoft published work on optimising agent "skill" artifacts, finding the optimised artifacts carry over both to different model sizes and across two competing agent harnesses. Why it matters: this is the first real evidence that the investment enterprises are making in agent scaffolding is portable rather than vendor-specific. A skill library becomes a durable asset that survives a model switch — which changes the procurement calculus from "pick a lab" to "pick a lab for now." Unverified Sources: MarkTechPost, The Neuron daily digest.

  5. Millennium and Anthropic are building a digital risk analyst on Claude. Anthropic published a joint case study with Millennium — one of the largest multi-strategy hedge funds in the world — describing an AI-powered digital risk analyst built on Claude. Why it matters: risk is the most conservative seat in the most conservative corner of finance. A named deployment there is worth more as a procurement unlock than a hundred pilots, and it lands in the same week regulators and banks are wiring up fraud-detection watch lists. Corroborated Sources: Anthropic/Claude blog, Capital Brief Standup.

Categorised News

Frontier & Big Tech

Meta ships Muse Code (Beta), a terminal coding agent on Muse Spark 1.2. Meta AI released a beta terminal-based coding agent backed by a new model, Muse Spark 1.2. It puts Meta directly into the Claude Code / Codex CLI category it had previously ceded, and pairs a coding-tuned model with its own harness rather than shipping weights alone. Unverified Source: MarkTechPost.

Gemini Robotics and Seedance 2.5 land in the same release window. Google pushed a Gemini Robotics update and ByteDance shipped Seedance 2.5 on the video side, both flagged in this week's model roundup without detailed specs. Treat as release signals, not benchmark claims. Unverified Source: AI Search weekly.

The first 48 hours of Claude Fable 5. Early-adopter writeups of Anthropic's Fable 5 circulated through the newsletter layer this week, focused on hands-on impressions rather than published evaluations. Worth watching for whether the reported behaviour holds up once third-party benchmarks land. Unverified Source: The AI Opportunities.

Apps / Dev Tools / Platforms

Marker v2: documents in, markdown out, three modes. Marker shipped v2 as a three-mode document-conversion pipeline producing markdown from arbitrary documents. Unglamorous and load-bearing: nearly every RAG and agent pipeline in production is bottlenecked on document parsing quality, not model quality. Unverified Source: MarkTechPost.

OpenAI's Codex Voice + browser + Sites workflow gets a public teardown. Nick Baumann of OpenAI walked through three advanced workflows on Lenny's How I AI: ChatGPT Voice as a logistics and travel assistant, building and deploying a live website from a single prompt, and automating UGC video editing from raw clips. Notable less for the tools than for what an insider actually does daily. Unverified Source: Lenny's Newsletter.

Wispr Flow adds a Notetaker. The dictation company extended into meeting capture, moving from "speak instead of type" toward the crowded ambient-notetaker category. Unverified Source: The Neuron.

Infrastructure & Ecosystem

NVIDIA releases Alpamayo 2 Super: a 34B open vision-language-action model for robotaxis, under OpenMDW-1.1. NVIDIA open-released a 34B VLA model targeting robotaxi and autonomous-driving stacks under the OpenMDW-1.1 licence. The strategic read is familiar: NVIDIA gives away the model layer to grow demand for the layer it sells. Unverified Source: MarkTechPost.

Agent security becomes its own vendor category. MCP server and LLM application security moved from conference-track curiosity to dedicated commercial programming this week, with vendor sessions specifically on securing AI agents and MCP servers. Inference The lag between "agents run in production" and "someone sells you agent security" has closed to roughly one quarter. Source: MarkTechPost newsletter.

Regions / Macro

EU opens a €30B AI gigafactory call: seven sites, 100K+ chips each. Brussels put out a call for seven AI gigafactory sites with a stated floor above 100,000 chips per site, backed by roughly €30B. Inference Read alongside the AI Act deadline slippage earlier this year, the European strategy has visibly shifted from regulating the technology to buying the capacity — subsidy over statute. Unverified Source: Creators' AI weekly digest.

South Korea signs a reported $950B in AI deals at a San Francisco summit. Korean industry and government agreements totalling roughly $950B were announced at an SF summit. The number is large enough that composition matters enormously — multi-year capex commitments and MOUs are not the same instrument, and the source does not break them out. Unverified Source: Creators' AI weekly digest.

Ken Griffin puts a number on losing access to Taiwan. The Citadel founder quantified what he considers the dominant tail risk in technology: US loss of access to Taiwanese semiconductor supply. Coming from someone who prices tail risk professionally, the framing carries more weight than the usual geopolitical commentary. Unverified Source: The AI Opportunities.

Market Cap / Valuation

SpaceX shares hold steady as a $100B insider lockup expires. Bloomberg reported SpaceX shares stayed flat through the expiry of a roughly $100B insider lockup — an unusually orderly outcome for an event of that size. Relevant to AI as a read on private-market depth for the mega-cap privates, a cohort that now includes several frontier labs. Corroborated Sources: Bloomberg, Capital Brief Standup.

Alex Karp's filter for which AI companies survive, and 19 new details on Cursor's rise. Two widely-circulated pieces this week: the Palantir CEO's framework for separating durable AI businesses from the rest, and a detailed reconstruction of Cursor's growth curve. Both are commentary rather than disclosure, but Cursor remains the cleanest case study in the AI application layer. Unverified Source: The AI Opportunities.

AI Gone Wrong / Disasters / Harms

CBA flags roughly $1B in suspected fraudulent transactions in an unusually complex scheme. Commonwealth Bank of Australia disclosed suspected fraudulent activity around the $1B mark, described as an unusually complex scheme. In parallel, Australian banks are building a new watch list to capture compromised staff, accountants, and lawyers. Inference The pairing with Millennium's Claude-based risk analyst is the story: financial-crime detection is becoming the first genuinely mandatory enterprise AI workload, because the alternative is a regulator. Corroborated Sources: Capital Brief Standup, Capital Brief.

Prompting Skill of the Week

Technique: Skill Extraction (writing the artifact, not the prompt). Best for: any task you have now run more than five times — weekly reporting, code review, RFP triage, incident writeups. This is the prompt-level version of what Microsoft's SkillOpt work formalises: stop tuning prompts, start producing a reusable skill artifact that survives model and harness changes.

  1. Pick one recurring task and find your single best past run — the output you'd happily show a client.
  2. Paste that transcript back to the model and ask it to reverse-engineer the procedure, not the answer: triggers, preconditions, steps, done-criteria, known failure modes.
  3. Strip everything model-specific. No "as Claude," no vendor-specific tool names, no phrasing that only works because of one model's quirks.
  4. Add explicit stop conditions and a definition of done. Most skill artifacts fail because they say what to do and never say when to stop.
  5. Test on a smaller, cheaper model. If the skill only works on the frontier model, you've encoded capability, not procedure.
  6. Iterate on the artifact file, not on the chat. Version it in git next to the code it supports.

Example prompt:

"Here is a transcript of a task I want to make repeatable. Reverse-engineer it into a reusable skill document with these sections: WHEN TO USE, INPUTS REQUIRED, STEPS (numbered, imperative), DONE WHEN, COMMON FAILURES. Do not include the specific answer from this transcript anywhere in the document. Then list the three places this skill would break if the input were different."

Common failure + fix: the skill silently memorises this week's answer, so it produces beautiful output on the example and nonsense on anything new. Fix: run it against three unseen inputs before you save it, and delete any line that would have to change for those three inputs to work.

New AI Tools

Marker v2. A three-mode document-to-markdown conversion pipeline. Audience: anyone building RAG or agent workflows over PDFs, scans, and office documents — which in practice means every internal AI project that has stalled on "the model can't read our files." Source: MarkTechPost.

Muse Code (Beta), Meta. A terminal coding agent powered by Meta's new Muse Spark 1.2 model, putting Meta into direct competition with Claude Code and Codex CLI. Audience: engineering teams already running terminal agents who want a third option, and anyone tracking whether Meta's model strategy has shifted from open weights toward its own harness. Source: MarkTechPost.

Wispr Flow Notetaker. Meeting capture from the dictation company, extending Flow from voice input into ambient note-taking. Audience: heavy meeting loads with an existing Flow habit; a crowded category where the differentiator is the rest of the workflow, not the transcription. Source: The Neuron.

AI Personality of the Week

Ken Griffin. The Citadel founder used the week to put an explicit number on the risk everyone in AI talks around: what happens to US technology if access to Taiwanese semiconductor supply is lost. It matters because of who is saying it — Griffin's business is pricing tail risk, and the AI industry's entire capex thesis assumes that particular tail stays folded. The timing is what makes it a story rather than a quote: the same week Griffin priced the chip risk, Washington moved on humanoid robot imports, Brussels opened a €30B chip-capacity call, and a peer multi-strategy fund in Millennium announced it is building a risk analyst on Claude. Finance is now simultaneously the loudest voice on AI supply-chain risk and one of the fastest adopters of the technology creating that risk. Source: The AI Opportunities.

Catch-All

"What if you're not supposed to have a long-term plan?" Lenny Rachitsky published an argument against the multi-year career and strategy plan — that in a fast-moving environment, optionality and rapid iteration beat commitment to a fixed destination. It circulated hard this week for an obvious reason: every roadmap written in 2026 has been invalidated at least twice, and the people writing them are looking for permission to stop. Inference For an executive audience the useful reframe is not "abandon planning" but "shorten the commitment horizon and lengthen the direction horizon" — keep the thesis, throw away the eighteen-month Gantt chart. Source: Lenny's Newsletter.

Show Notes (bullets only)

  • OpenAI reportedly at 1 billion users, with 99.8% of tokens now served inside agentic loops rather than chat.
  • Moonshot publishes full Kimi K3 weights; DeepSeek ships V4-Flash-0731 on Hugging Face under an MIT licence.
  • MiniMax H3 and new AMD models land the same week — four open releases in seven days.
  • The US moves to ban foreign-made humanoid robots, explicitly targeting China.
  • Microsoft's SkillOpt shows optimised agent skills transfer across model sizes and between Codex and Claude Code.
  • Millennium and Anthropic announce a digital risk analyst built on Claude.
  • NVIDIA open-releases Alpamayo 2 Super, a 34B vision-language-action model for robotaxis, under OpenMDW-1.1.
  • EU opens a €30B call for seven AI gigafactories, 100K+ chips each.
  • South Korea signs a reported $950B in AI deals at a San Francisco summit.
  • Ken Griffin puts a number on the US losing access to Taiwanese chips.
  • Meta ships Muse Code (Beta), a terminal coding agent on Muse Spark 1.2.
  • CBA flags ~$1B in suspected fraud; Australian banks build a watch list for compromised staff, accountants, and lawyers.

Weekly Patterns (Inference)

  1. Inference Open weights are now a Chinese-led category. Kimi K3 full weights and an MIT-licensed DeepSeek release in the same week means the default self-hosted option for a Western enterprise is increasingly a Chinese model — a procurement and security conversation most boards have not had yet.

  2. Inference The moat moved from weights to skills, and then the skills turned out to be portable too. SkillOpt's cross-harness transfer result undercuts the main lock-in story the agent vendors have been telling. If the artifact survives a switch from Codex to Claude Code, switching cost collapses.

  3. Inference Software commoditises, hardware nationalises. Free weights, portable skills, and open driving models on one side; humanoid import bans, €30B gigafactory calls, and $950B bilateral packages on the other. Value is migrating to whatever cannot be downloaded.

  4. Inference "Agentic" is now the default consumption mode, and nobody's finance function is ready. If OpenAI's 99.8% figure is even roughly right, per-seat licensing is already the wrong unit — but note the number is self-reported and the definition of "agentic token" is doing enormous work.

  5. Inference Finance is the beachhead for enterprise agents, and the entry point is risk, not revenue. Millennium's risk analyst, CBA's $1B fraud disclosure, and the new banking watch lists are one story: the first mandatory AI workload is the one where the alternative is a regulator.

  6. Inference Model releases have stopped being events. Four significant open releases plus Gemini Robotics, Seedance 2.5, and Meta's Muse Spark 1.2 landed in a single week with no individual release dominating. When a frontier drop is a bullet point, the differentiation has moved somewhere else — distribution, harness, and data.

  7. Inference Pre-emptive controls are the new pattern. Chip export restrictions arrived years into the boom; the humanoid robot restriction arrives before the category has commercial scale. Expect the same forward-leaning posture on agent infrastructure and inference capacity next.

// Deep dives from this episode