All the AI signal. None of the hype.

A free daily AI newsletter: everything that actually happened in AI today. News, research, tools, and what to build.

Read today's issue → Get the free Top 100 AI tools list →
Tracking primary sources OpenAIAnthropicGoogle DeepMindMeta AIarXivHugging FaceSEC filings
The day's signal, from primary sources.
News
Oct 6 · News

Reflection ships Beam, a 501B open-weight model aimed at China's frontier labs

Reflection AI, the Nvidia-backed startup that's raised roughly $4.7B total, shipped Beam yesterday: 501B total parameters, 23B active in a MoE, trained on 23.8T tokens, context stretched to 1M. The pitch is parity with Z.ai's GLM-5.2 at 3 to 4x less inference compute, a claim TechCrunch notes hasn't been independently verified.

Read at reflection.ai →
Models
Oct 6 · Models

Anthropic

Claude Opus 5.5 and Sonnet 5.5 are now GA in AWS GovCloud, hooked into Claude Code for regulated and ITAR work. Agencies that couldn't touch Claude for compliance reasons just lost that excuse.

Read at aws.amazon.com →
Models
Oct 6 · Models

OpenAI

ChatGPT is getting a visual ad format, with attribution partners and brand-suitability tools bolted on for advertisers. The thing that answers your question is about to also pitch you a product.

Read at openai.com →
Policy
Oct 6 · Policy

OpenAI spelled out how it's handling the EU AI Act's text-provenance mandate: watermark generated text, build a detector for it, then hand that detector to researchers first, not the public.

The compliance approach binds OpenAI under the Act's content-marking rule, and the staggered rollout is the real tell. They're not ready to let outsiders stress-test the detector yet.

Read at openai.com →
News
Oct 5 · News

Supabase buys Turso, bets $150M that the next database customer is an agent

Supabase announced it's acquiring Turso on October 2 for an undisclosed sum, with founder Glauber Costa joining as Head of Agentic Services. The deal lands alongside a $150M round led by GIC, with CapitalG, IronArc, and SquarePeg, just four months after Supabase's $500M Series F at a $10.5B valuation, also GIC-led.

Read at supabase.com →
News
Oct 4 · News

Argo-Bench just showed the best data agent on the market clears a third of realistic tasks

Argo-Bench drops data agents into a simulated NYC delivery warehouse: 235 tables, 7.5 billion rows, 210 tasks that end in real actions like banning an account or allocating a budget, not a SQL-grading rubric. The best of 14 frontier and open-weight models broke 95 points on just 34.8% of those tasks, averaging 59.5 overall.

Read at arxiv.org →
New AI tools worth trying. Trending projects and launches.
a

amitshekhariitbhu/ai-system-design

AI System Design - Learn how to design AI systems built on LLMs, RAG, and AI Agents step by step. Markdown. ★636 on GitHub. Free, open source.

Visit · github.com →
X

XHToken/Spark-X2.5

Spark-x2.5 open model series. Pushing the Limits of Agentic Capabilities in On-Device Models ★660 on GitHub. Free, open source.

Visit · github.com →
P

Ptero-ai/Ptero

AI Chat with powerful models for free PHP. ★1,465 on GitHub. Free, open source.

Visit · github.com →
A

AWS's Adjudicated Query pattern

bounds a chat agent to six typed operations over a deterministic rules engine, so compliance answers ship with a "completeness receipt": compliant plus in-breach plus ambiguous plus unreadable has to equal the number scanned. Same instinct as today's lead: don't let the model grade its own homework.

Visit · aws.amazon.com →
S

ServiceNow's AutoSynthData hunts for an agent's actual failure modes and generates synthetic tasks to train against them.

On ServiceNow's own ITSM benchmark that took Pass@1 from 18.77% to 27.18%. A way to go looking for the state-corrupting cases before a benchmark finds them for you.

Visit · huggingface.co →
d

devagrawal09/jev-review

A staged code-review workflow and local dashboard built with TypeSafe Jev. TypeScript. ★661 on GitHub. Free, open source.

Visit · github.com →
What to build next. Real gaps with the receipts: the evidence, why it's tractable now, and where to start.
weekend buildAI agents / reliability

Agents need a recovery layer, not a restart button

A single failed tool call can sink an entire multi-step agent run. The opportunity is a small checkpoint-and-resume layer that lets builders recover from the last good state.

StartWrap each agent step in a checkpoint. On failure, replay from the last clean state with a local repair instruction instead of restarting the run.
weekend buildAI agents / memory

Agent memory needs expiration dates

Agent memory systems are good at saving facts and bad at knowing when those facts have expired. Stale memory is a quiet product bug waiting to happen.

StartStore every memory with timestamp, confidence, and source. When new context contradicts it, supersede the old fact instead of appending another one.
startup-scaleDev tools / code review

AI code needs a missing-requirements reviewer

AI can write code faster than teams can review it. The gap is not syntax; it is the requirements the prompt never mentioned.

StartRun a review pass on each AI-authored PR that checks auth, input validation, secrets, permissions, data retention, and threat-model omissions.