Agent 107 briefings A $400K Open Image Stack, and Distillation as a Catalyst An AI Scientist Reproduces a GWAS Inside a Hospital Firewall The Best Model Gets Idea Lineage Right Only 27% of the Time View topic →
Multimodal 105 briefings A $400K Open Image Stack, and Distillation as a Catalyst A 3B-Active MoE Matches Last Year's 32B An AI Scientist Reproduces a GWAS Inside a Hospital Firewall View topic →
Evaluation 97 briefings A $400K Open Image Stack, and Distillation as a Catalyst An AI Scientist Reproduces a GWAS Inside a Hospital Firewall Multi-Turn Pressure Pushes CLI Agents to 100% Compliance View topic →
Image Gen 88 briefings A $400K Open Image Stack, and Distillation as a Catalyst A 3B-Active MoE Matches Last Year's 32B The Best Model Gets Idea Lineage Right Only 27% of the Time View topic →
Training 86 briefings A $400K Open Image Stack, and Distillation as a Catalyst Structural Reasoning Takes 67 SOTAs, Async RL Ships in GLM-5.2 RL Reaches Image Generation, Gemma 4 Goes Open View topic →
Safety 80 briefings RL Reaches Image Generation, Gemma 4 Goes Open Verification as the Fourth Scaling Axis Avatar Resolution Doubles, Latency Holds at 200ms View topic →
AI for Science 70 briefings RL Reaches Image Generation, Gemma 4 Goes Open Avatar Resolution Doubles, Latency Holds at 200ms Latent-Space Scoring Cuts Video Generation to 1-4 Steps View topic →
Efficiency 69 briefings The Model as Its Own Teacher, and Why Splitting Crowds Wins RL Reaches Image Generation, Gemma 4 Goes Open Verification as the Fourth Scaling Axis View topic →
Architecture 67 briefings The Model as Its Own Teacher, and Why Splitting Crowds Wins The Best Model Gets Idea Lineage Right Only 27% of the Time Latent-Space Scoring Cuts Video Generation to 1-4 Steps View topic →
Robotics 65 briefings A 3B-Active MoE Matches Last Year's 32B An AI Scientist Reproduces a GWAS Inside a Hospital Firewall Structural Reasoning Takes 67 SOTAs, Async RL Ships in GLM-5.2 View topic →
Video Gen 64 briefings An AI Scientist Reproduces a GWAS Inside a Hospital Firewall The Best Model Gets Idea Lineage Right Only 27% of the Time RL Reaches Image Generation, Gemma 4 Goes Open View topic →
Reasoning 56 briefings Memory Makes Agents Sycophantic; Visual Reasoning Hits 93.2% Frontier Agents Finish One Task in Five at 1.6-Hour Length Knowing When to Stop Doubles an Agent's Recall View topic →
Retrieval 53 briefings Verification as the Fourth Scaling Axis Memory Makes Agents Sycophantic; Visual Reasoning Hits 93.2% A 35B Agent Reaches for Trillion-Scale, and Async Lag Is Overrated View topic →
Interpretability 51 briefings RL Reaches Image Generation, Gemma 4 Goes Open Avatar Resolution Doubles, Latency Holds at 200ms 0.6B Matches 32B, 50x Less VRAM View topic →
Code Intelligence 34 briefings The Model as Its Own Teacher, and Why Splitting Crowds Wins Leaderboards Don't Predict Deployment; Robot Arm Self-Trains to 99% Two Loops Take SWE-bench From 43 to 64 View topic →