Good morning, {{first_name | AI enthusiast}}.
Fermat’s Last Theorem stumped mathematicians for 358 years until Andrew Wiles finally cracked it by hand in 1995. Now Anthropic says Claude turned that proof into a fully computer-checkable formalization — 13 million lines of code, built mostly on its own in 11 days.
Mathematicians are calling the result “robust enough to be built upon,” but does an AI verifying a human proof count as genuine reasoning, or just extremely patient bookkeeping? Two other labs spent the week testing similar questions about how far autonomous AI can run without supervision.
Today in AI Brief:
Claude formalizes Fermat’s Last Theorem
Meta’s AI research agent beats 4,000 teams
Google’s Lyria 3.5 writes full songs
Build a Holiday Creator Affiliate Program in 90 Days
Creators lock in holiday content calendars 90 days out, before brands figure out commissions. Waiting too long to launch an affiliate program means less runway to build demand and a missed shot at the best partnerships.
The 90-Day Holiday Sprint covers commissions, recruiting, and scaling a program at Day 30, 60, and 90.
Claude Formalizes Fermat’s Last Theorem
In Brief: Anthropic says Claude produced the first fully computer-verified formalization of Andrew Wiles’ 1995 proof of Fermat’s Last Theorem, converting the 129-page proof into machine-checkable Lean code over 11 days, largely without human help.
The Details:
Claude generated 13 million lines of Lean code — the largest formal proof file ever produced — proving 29,500 intermediate theorems along the way and consuming roughly 6 billion output tokens across several dozen agents.
The run only became possible after Anthropic gave Claude access to Prove2Me, an open-source tool built at Columbia that helps agents pick the most useful next step in a long formalization chain.
Mathematician Kevin Buzzard, who leads Imperial College’s formal-proof project, said the result shows AI autoformalization is now “robust enough to be built upon” across algebra, geometry, and number theory.
Take Away:
Checking a landmark proof line-by-line is a different, harder bar than the benchmarks AI models usually chase. If autoformalization keeps compressing years of manual verification into days, entire fields of pure math could gain a machine-checkable backbone within a few years.
Meta’s AIRA₃ Wins Kaggle Gold Against 4,000 Teams
In Brief: Meta’s autonomous research system AIRA₃ placed 8th out of roughly 4,000 teams in a live, NVIDIA-run Kaggle competition to fine-tune a 30B-parameter Nemotron model for better reasoning, earning a gold medal against human data scientists using the same public tools.
The Details:
AIRA₃’s agents worked in separate environments, posting ideas to a shared forum and swapping code through shared files instead of routing through a central manager — letting agents build on each other’s experiments as they went.
The system pushed the model’s benchmark rank from a 72.7% baseline to 81.5% within 24 hours, then to 83.1% after 72 hours — a margin Meta says beat prior research agents by roughly 9 percentage points.
Meta says the win signals AIRA₃ can improve a targeted AI capability at a level similar to human experts, though the system itself hasn’t been released outside the company.
Take Away:
A research agent outscoring thousands of human competitors, using tools everyone had equal access to, is a rare concrete data point for “AI improving AI” claims that are usually hard to verify. It’s also a preview of labs scaling the research process itself, not just the models that research produces.
SPONSORED BY CAELITH.AI
Stop Paying for 10 Tools. One AI Does It All.
Most e-commerce sellers are running their store across 6 to 10 separate tools — and spending more time managing software than growing their business. StoreClaw replaces your entire stack with one autonomous AI engine that monitors competitors, optimizes listings, automates marketing, and tracks real profit across Shopify, Amazon, and beyond.
It doesn't wait for you to ask. It runs 24/7 in the background, so you wake up to a full dashboard instead of a list of things you forgot to check.
Connect your store, and StoreClaw gets to work — no prompts, no complex setup, no six-app stack.
Free to start. No credit card required.
Google’s Lyria 3.5 Writes Full Songs Inside Gemini
In Brief: Google rolled out Lyria 3.5, its latest music-generation model, directly inside the Gemini app, AI Studio, and API — letting anyone type a prompt and get a finished, up to three-minute song with vocals and full arrangement.
The Details:
The model outputs high-fidelity 44.1kHz stereo audio and lets users pick a genre, choose vocal or instrumental styles, or start from templates for things like background music or a personalized birthday song.
Lyria 3.5 first launched inside Google’s separate Flow Music tool in July, and this wider Gemini rollout adds improvements to musicality and lyric generation, according to reporting on the release.
Google says the model was trained on licensed content, a contrast it’s drawing directly with rivals like Suno, though it hasn’t detailed which catalogs.
Take Away:
Creative AI tools win attention fastest when they’re easy to try, and a full song generator built into an app hundreds of millions of people already open lowers that bar to almost nothing. Expect the “licensed training data” pitch to keep showing up as music-AI copyright lawsuits pile up elsewhere.
Everything else in AI
Crusoe raised $3B at a $30B valuation, days after signing a $13B, five-year cloud contract to supply trading firm Jane Street with GPUs and AI infrastructure.
GitHub launched Project HydraFusion, a Copilot CLI research preview that routes coding tasks across multiple AI models and matches Claude Opus 5’s quality at up to 67% lower cost.
OpenAI says it has built an “automated research intern” that can handle multi-day research tasks under human direction, a step toward a fully autonomous AI researcher targeted for March 2028.
Hikers were rescued from Mount Shasta after Google’s Gemini reportedly told them to pack far less food and water than their overnight summit attempt required.


