// Article · July 24, 2026 · 2 min read
Devices & Robotics — W30: Samsung makes on-device AI the foldable's pitch, and the run-it-local model tier fills in
A quiet week for robots, a structural one for edge inference — the models small enough to run on your own hardware are finally arriving in a crop.
Most of this week's AI oxygen went to a 2.8-trillion-parameter cloud model, a lawsuit, and a chipmaker's earnings. The physical-world thread underneath all of it was on-device inference — the models small enough to run at the edge, and the hardware built to sell them as a feature. It's a positioning week, not a benchmark week, but the direction is clear.
Samsung makes on-device AI the foldable's headline spec
Samsung unveiled its next-generation Galaxy foldables this week and leaned on AI features as the differentiator — not the hinge, not the crease, the model. The read for anyone tracking edge inference: the phone makers have decided the on-device assistant is the marketing. The folding form factor has quietly demoted itself from the story to the vehicle for the story. What's still missing is the spec sheet that matters — which model, running how fast, on what silicon — but the framing shift is the tell. Via The Neuron.
Cisco open-sources two security models small enough to run in CI
Cisco's Foundation AI group released Antares-350M and Antares-1B — open-weight small language models purpose-built to hunt and localize vulnerabilities in source code. At 350M and 1B parameters, these are sizes you run on every commit, self-hosted, with nothing leaving the building. The pitch is economics and data control, not raw capability: a 1B model in your CI pipeline is a fundamentally different security posture than shipping your codebase to a frontier API. It's the clearest sign this week that the specialized on-prem tier is where small models win first. Via MarkTechPost.
The run-it-local model tier keeps filling in
Buried under Kimi K3's cloud-scale release, the same week brought Bonsai 27B and OpenAI's Codex Micro — sizes that sit squarely in the run-it-yourself range. A 27B model, quantized, fits on a single high-memory consumer GPU; a "micro" coding model is built for exactly the latency and privacy budget the cloud can't hit. These are still lightly sourced, so treat the parameter counts as reported rather than benchmarked. But the pattern holds across the week: for every headline giant, a device-class model now ships alongside it, almost as an afterthought. Via AI Search.
What to watch next week: whether Samsung backs the "AI foldable" framing with an actual on-device number — tokens per second, which model, on which NPU — or leaves it at vibes, and whether any of this quarter's small models get a real edge-hardware demo instead of a HuggingFace card. Right now the physical-world AI story is all positioning. The next quarter should force it to show its timings.
This post is also published on our Substack newsletter at edge-ai.forum. Subscribe for the weekly roundup direct to your inbox — fresh AI news, executive context, and devices + robotics every Friday morning.
// Related
July 24, 2026 · 3 min
Executive Roundup — W30: The week leverage moved from models to compute, courts, and the open-weight line
July 24, 2026 · 2 min
LLM Weekly — W30: Kimi K3 makes the open frontier 2.8 trillion parameters — and Chinese
August 14, 2026 · 3 min
Devices & Robotics — W33: Dyna-2 learns from human video, and Xiaomi decides it wants to build robots