← All Series

FRONTIER DIGEST // WEEKLY AI UPDATES

AI Week in Review.

A plain-English briefing on what changed in AI this week: new models and products, company news, policy and safety, useful tools, and the research that matters. Each edition connects the stories and puts sources beside the entries. Announcements, vendor or author measurements, and editorial interpretation are kept distinct. Coverage dates appear on every edition.

12 editions
10–15 min to catch up
weekly new briefing
source-linked news + research

EARLIER EDITIONS

Previous weeks · newest to oldest
edition #10 · Aug 30, 2026

The Cost of Putting AI to Work. live

Jalapeño's first results, efficient open models, computer-use agents, NVIDIA's earnings, and all ten papers from the weekly research selection.

edition #9 · Aug 23, 2026

The Harness Started Training the Model. live

OpenAI paces Astra, Claude designs protein binders tested by outside labs, and ten papers explore the harness as a training-time and safety-time object.

edition #8 · Aug 16, 2026

The State Your AI Keeps Is the System You Ship. live

NVIDIA’s $500B infrastructure-financing move, Meta’s personal-superintelligence direction, active EU transparency duties, and ten papers point to the same practical conclusion: capacity, access, compliance, and persistent system state now need to be engineered together.

edition #7 · Aug 9, 2026

Models Compete. Runtimes Decide. live

Astra's reported cyber-driven slowdown and a private U.S. review framework meet ten papers that make the same systems argument: agent capability belongs to the model–runtime pair. Failure ownership, executable harness repair, memory, matched-budget reasoning, long-horizon coherence, and serving efficiency all move above the checkpoint layer.

edition #6 · Aug 2, 2026

The week the sandbox became the security perimeter live

Anthropic's three real-world cyber-evaluation incidents make containment the lead story; NVIDIA launches the Open Secure AI Alliance; Microsoft introduces MAI-Cyber-1-Flash and Project Perception; Kimi K3's full weights and technical report arrive; and ten papers connect agent interfaces, role drift, invisible reasoning, context and filesystem memory, agentic RL, TPU kernels, higher-order optimizers, and self-speculating tool calls.

edition #5 · Jul 26, 2026

The week the harness became the security boundary live

OpenAI's escaped cyber evaluation turns the harness into the week's security boundary. Gemini Flash, Presence, Health, AMD's Anthropic compute partnership, and the open-weights policy fight map the product, money, and policy layer; ten papers connect behavior maps, executable skills, programmatic memory, hidden workspaces, completeness, structured-output drift, memory attacks, 2D text, and robot fast weights.

edition #4 · Jul 19, 2026

The week agents learned to watch themselves live

Kimi K3 and Thinking Machines' Inkling split the open-weight frontier between scale and customizability, while WANDR exposes the limits of evidence-heavy research agents. Ten papers connect self-improvement, metacognition, routing, matched-budget harness evaluation, failure tracing, filtered sabotage monitoring, distribution-matching RL, interactive worlds, and cross-embodiment robotics.

edition #3 · Jul 12, 2026

The week the harness became the product live

Three flagships on one Thursday — GPT-5.6 goes GA (Sol/Terra/Luna), SpaceXAI's "Opus-class" Grok 4.5 at $2/$6, and Meta starts charging with Muse Spark 1.1 — while the research converges on the orchestration layer: The Harness Effect's 41% cost cut with models held constant, Verification as a Scaling Axis, ReContext, Oxford's failure taxonomy (scaffolding ≠ robustness), Always-On Agents' durable-state survey, HOLA, and Puzzle-75B.

edition #2 · Jul 5, 2026

The week verification became the whole game live

The whole week in AI: the frontier's center of gravity moves (Claude Sonnet 5 ships as the agentic free default; Anthropic reportedly past OpenAI; Fable 5 returns), the money into physical/agentic AI, policy, and the ten research papers threaded around one throughline — the evaluator is the load-bearing part and freezing it kills progress (Red Queen Gödel Machine, Verification Horizon, RLMF, Paper Assistant Tool), with self-evolving code reaching robots and chips (ASPIRE, HORIZON).

edition #1 · Jun 25, 2026

The week agents became systems (and evaluation grew teeth) live

The whole week in AI: the lead story (GPT-5.6's system card going "High" on cyber + biology), models & releases (Sakana Fugu, GLM-5.2 open weights), business & funding (the inference megarounds), policy & safety, the ten research papers, tools & open source, and a ranked "what matters most."


WHAT YOU GET EVERY WEEK

  • The biggest change. One lead story explains the development that shaped the week.
  • What shipped. New models, products, tools, and open-source releases in plain language.
  • The wider picture. Company news, funding, policy, and safety developments that affect builders.
  • Research worth your time. The useful papers, grouped by theme with a short explanation of why they matter.
  • A clear takeaway. A ranked “what matters most” list and the patterns connecting the week.