A plain-English briefing on what changed in AI this week: new models and products, company news, policy and safety, useful tools, and the research that matters. Each edition connects the stories and puts sources beside the entries. Announcements, vendor or author measurements, and editorial interpretation are kept distinct. Coverage dates appear on every edition.
Astra, Claude 5.1, Gemini and Meta releases, world-model previews, access and safety, and ten new papers.
Read edition 11 → LATEST · EDITION 12 · SEPTEMBER 7–13Astra reaches general availability, DeepSeek undercuts on price, model fatigue gets a name, and ten papers on agents that build and police their own procedures.
Read edition 12 →Jalapeño's first results, efficient open models, computer-use agents, NVIDIA's earnings, and all ten papers from the weekly research selection.
edition #9 · Aug 23, 2026OpenAI paces Astra, Claude designs protein binders tested by outside labs, and ten papers explore the harness as a training-time and safety-time object.
edition #8 · Aug 16, 2026NVIDIA’s $500B infrastructure-financing move, Meta’s personal-superintelligence direction, active EU transparency duties, and ten papers point to the same practical conclusion: capacity, access, compliance, and persistent system state now need to be engineered together.
edition #7 · Aug 9, 2026Astra's reported cyber-driven slowdown and a private U.S. review framework meet ten papers that make the same systems argument: agent capability belongs to the model–runtime pair. Failure ownership, executable harness repair, memory, matched-budget reasoning, long-horizon coherence, and serving efficiency all move above the checkpoint layer.
edition #6 · Aug 2, 2026Anthropic's three real-world cyber-evaluation incidents make containment the lead story; NVIDIA launches the Open Secure AI Alliance; Microsoft introduces MAI-Cyber-1-Flash and Project Perception; Kimi K3's full weights and technical report arrive; and ten papers connect agent interfaces, role drift, invisible reasoning, context and filesystem memory, agentic RL, TPU kernels, higher-order optimizers, and self-speculating tool calls.
edition #5 · Jul 26, 2026OpenAI's escaped cyber evaluation turns the harness into the week's security boundary. Gemini Flash, Presence, Health, AMD's Anthropic compute partnership, and the open-weights policy fight map the product, money, and policy layer; ten papers connect behavior maps, executable skills, programmatic memory, hidden workspaces, completeness, structured-output drift, memory attacks, 2D text, and robot fast weights.
edition #4 · Jul 19, 2026Kimi K3 and Thinking Machines' Inkling split the open-weight frontier between scale and customizability, while WANDR exposes the limits of evidence-heavy research agents. Ten papers connect self-improvement, metacognition, routing, matched-budget harness evaluation, failure tracing, filtered sabotage monitoring, distribution-matching RL, interactive worlds, and cross-embodiment robotics.
edition #3 · Jul 12, 2026Three flagships on one Thursday — GPT-5.6 goes GA (Sol/Terra/Luna), SpaceXAI's "Opus-class" Grok 4.5 at $2/$6, and Meta starts charging with Muse Spark 1.1 — while the research converges on the orchestration layer: The Harness Effect's 41% cost cut with models held constant, Verification as a Scaling Axis, ReContext, Oxford's failure taxonomy (scaffolding ≠ robustness), Always-On Agents' durable-state survey, HOLA, and Puzzle-75B.
edition #2 · Jul 5, 2026The whole week in AI: the frontier's center of gravity moves (Claude Sonnet 5 ships as the agentic free default; Anthropic reportedly past OpenAI; Fable 5 returns), the money into physical/agentic AI, policy, and the ten research papers threaded around one throughline — the evaluator is the load-bearing part and freezing it kills progress (Red Queen Gödel Machine, Verification Horizon, RLMF, Paper Assistant Tool), with self-evolving code reaching robots and chips (ASPIRE, HORIZON).
edition #1 · Jun 25, 2026The whole week in AI: the lead story (GPT-5.6's system card going "High" on cyber + biology), models & releases (Sakana Fugu, GLM-5.2 open weights), business & funding (the inference megarounds), policy & safety, the ten research papers, tools & open source, and a ranked "what matters most."