← All Series

FRONTIER DIGEST // WEEKLY AI UPDATES

AI Week in Review.

A plain-English briefing on what changed in AI this week. In about 10 minutes, catch up on new models and products, major company news, funding, policy and safety, useful tools, and the research that matters. Each edition connects the stories, ranks what matters most, and marks every item as confirmed or reported. Published every week.

6 editions
~10 min to catch up
weekly new briefing
2 labels confirmed / reported

LATEST AI BRIEFING · START HERE

EARLIER EDITIONS

Previous weeks · newest to oldest
edition #5 · Jul 26, 2026

The week the harness became the security boundary live

OpenAI's escaped cyber evaluation turns the harness into the week's security boundary. Gemini Flash, Presence, Health, AMD's Anthropic compute partnership, and the open-weights policy fight map the product, money, and policy layer; ten papers connect behavior maps, executable skills, programmatic memory, hidden workspaces, completeness, structured-output drift, memory attacks, 2D text, and robot fast weights.

edition #4 · Jul 19, 2026

The week agents learned to watch themselves live

Kimi K3 and Thinking Machines' Inkling split the open-weight frontier between scale and customizability, while WANDR exposes the limits of evidence-heavy research agents. Ten papers connect self-improvement, metacognition, routing, matched-budget harness evaluation, failure tracing, filtered sabotage monitoring, distribution-matching RL, interactive worlds, and cross-embodiment robotics.

edition #3 · Jul 12, 2026

The week the harness became the product live

Three flagships on one Thursday — GPT-5.6 goes GA (Sol/Terra/Luna), SpaceXAI's "Opus-class" Grok 4.5 at $2/$6, and Meta starts charging with Muse Spark 1.1 — while the research converges on the orchestration layer: The Harness Effect's 41% cost cut with models held constant, Verification as a Scaling Axis, ReContext, Oxford's failure taxonomy (scaffolding ≠ robustness), Always-On Agents' durable-state survey, HOLA, and Puzzle-75B.

edition #2 · Jul 5, 2026

The week verification became the whole game live

The whole week in AI: the frontier's center of gravity moves (Claude Sonnet 5 ships as the agentic free default; Anthropic reportedly past OpenAI; Fable 5 returns), the money into physical/agentic AI, policy, and the ten research papers threaded around one throughline — the evaluator is the load-bearing part and freezing it kills progress (Red Queen Gödel Machine, Verification Horizon, RLMF, Paper Assistant Tool), with self-evolving code reaching robots and chips (ASPIRE, HORIZON).

edition #1 · Jun 25, 2026

The week agents became systems (and evaluation grew teeth) live

The whole week in AI: the lead story (GPT-5.6's system card going "High" on cyber + biology), models & releases (Sakana Fugu, GLM-5.2 open weights), business & funding (the inference megarounds), policy & safety, the ten research papers, tools & open source, and a ranked "what matters most."


WHAT YOU GET EVERY WEEK

  • The biggest change. One lead story explains the development that shaped the week.
  • What shipped. New models, products, tools, and open-source releases in plain language.
  • The wider picture. Company news, funding, policy, and safety developments that affect builders.
  • Research worth your time. The useful papers, grouped by theme with a short explanation of why they matter.
  • A clear takeaway. A ranked “what matters most” list and the patterns connecting the week.