Research — AI news
New AI research in plain language: models, benchmarks, training methods and results that move the field.
Research
A new framework maps AI agency risk from app to chip level
A new arXiv framework scores AI agency delegation by stack layer, from app permissions to chip-level control.
Research
Self-refinement pipelines need separate compute budgets, study finds
New research quantifies how splitting compute between generation and critique changes self-refinement outcomes.
Research
Nobody will confirm who built Ox Alpha, and that's the point
An unclaimed model called Ox Alpha is topping benchmarks, and nobody will say who built it.
Research
AI is making scientists do more work, not better work, study finds
A new study complicates the AI-productivity story in research, with real stakes for R&D ROI calls.
Research
Inherent says its AI beat Anthropic and OpenAI at replicating research
A DeepMind-alumni startup claims its agent outperforms rivals at reproducing scientific results.
Research
AI world models fail when they ignore what people believe
New research finds that predicting human actions requires modeling what they believe, not just what they do.
Research
Open-weight models are catching closed ones twice as fast
SemiAnalysis finds the open-closed AI gap now halves every model era, not just every year.
Research
A third of the web's new pages are now AI-written, study finds
A new study finds roughly a third of pages published since ChatGPT's launch show AI authorship.
Research
DeepSeek V4 Flash brings near-Opus vision at a fraction of the cost
DeepSeek's cheap multimodal model puts pricing pressure on OpenAI and Anthropic's flagships.
Research
DeepMind traces 15 years of game AI, from Atari to EVE Online
How 15 years of game-playing agents built the RL toolkit behind today's LLM agents
Research
How researchers are teaching AI agents tasks from raw usage logs
A new method skips manual annotation and RLHF by inducing task models directly from computer-use logs.
Research
Why AI detectors still work: it's the guardrails, not the writing
Post-training safety tuning, not raw language ability, is what leaves LLM text with a detectable signature.
Research
Richard Sutton says synthetic data is AI's big mistake
The reinforcement learning pioneer warns that synthetic data can't capture the real world's complexity.
Research
China's frontier AI models have caught up with the West
Frontier Radar's fourth report finds the once-wide US-China AI gap has narrowed to nearly nothing.
Research
Generalist AI's GEN-1.5 teaches robots new tasks from one demo
One-shot imitation learning could let robotics teams skip massive labeled datasets and ship faster.
Research
Terence Tao: AI proofs risk math's biggest crisis since Gödel
The Fields medalist says fluent but unverifiable AI proofs threaten mathematics' peer-review system.
Research
Anthropic is keeping its most powerful model entirely in-house
Anthropic's most capable model runs only inside the company, with no public API or Claude.ai access.
Research
New study tracks how pretraining can teach then erase a single fact
A new arXiv paper tracks a single fact through pretraining — learned midway, then measurably lost.
Research
A new attention layer gives AI models free uncertainty estimates
Lévy Attention swaps softmax for a stochastic operator that reports its own confidence for free.
Research
Researchers fine-tune AI to search sound libraries by humming
A new arXiv paper zeroes in on how finetuning choices shape AI systems that retrieve sounds from vocal imitations.
Research
A new paper rethinks distillation for long-context LLMs
An arXiv paper proposes group-calibrated, on-policy distillation to fix long-context model compression.
Research
Why robot dexterity training is starting to look like LLM training
A new arXiv paper named ADEPT applies a pretrain-then-RL-post-train recipe to dexterous robot hands.
Research
SPADE trains AI agents via self-play in synthetic code environments
SPADE lets AI agents write, run, and solve their own coding challenges — no human-labeled tasks required.
Research
Why AI still hasn't won the public over, four years in
Despite billions in investment and near-daily model upgrades, public trust in AI has barely moved.
Research
VentureBeat adds a Lead Analyst role to steer its enterprise AI coverage
Rob Strechay becomes VentureBeat's first Lead Analyst, folding research-firm-style output into its AI journalism.
Research
AI's self-improvement problem is becoming an engineering question
MIT Tech Review's August 19 roundup reframes recursive self-improvement from sci-fi risk to a build problem.
Research
Anthropic's $2 trillion IPO talk is the AI boom's biggest test
A potential $2 trillion valuation would force Anthropic to prove AI spending finally pays off.
Research
Why GLM-5.3's benchmark score deserves a second look
Zhipu's flagship model tops charts, but the real signal sits in the fine print, not the leaderboard rank.
Research
The AI usage numbers everyone quotes, nobody can verify
Two landmark studies tried to map how people actually use AI chatbots — both left the core question open.
Research
Cutting the data needed to predict how materials block sound
A new arXiv preprint blends physics-based priors with machine learning to cut data needs for acoustic prediction.
Research
AutoSR reframes symbolic regression as a search over research states
A new AutoSR method treats equation discovery as navigating states in a research process, not brute-force search.
Research
New spectral-gap bounds sharpen two convex-sampling algorithms
A new arXiv paper analyzes how fast Hit-and-Run and Coordinate Hit-and-Run samplers actually converge.
Research
AlphaE: a fresh run at matrix multiplication's speed limit
A new arXiv paper applies modern optimization and an Alpha-style search to shrink ω.
Research
Inverse reinforcement learning gets a Q-function makeover on arXiv
A new preprint merges Q-learning with variational inference to make reward inference from demos more efficient.
Research
China is betting its data can power the world's AI models
Beijing is marketing its vast surveillance-era data trove as the raw material the world's AI industry needs.
Research
Why decoupling evidence from decisions could speed up AI training
A new arXiv preprint separates evidence gathering from decision aggregation to cut redundant compute in AI pipelines.
Research
A new method speeds up signal detection in massive MIMO systems
A new arXiv paper targets faster signal detection in massive-antenna MIMO networks, with lessons for AI compute design.
Research
Moral neutrality in AI is a choice, not a default setting
A new arXiv paper argues that claiming AI moral neutrality is itself a design decision developers must own.
Research
A new method lets AI models carry training state across sessions
A new arXiv paper proposes carrying model training state across sessions instead of restarting cold.
Research
Marionette helps AI agents predict and visualize world states
A new arXiv framework called Marionette forecasts world states and renders object geometry for AI agents.
Research
Amazon will train AI models on Twitch streamers' content
Amazon plans to train AI models on Twitch streams, turning live video and chat into a new data source.
Research
Why mathematicians trust LLMs with numbers but not new ideas
Leading mathematicians say LLMs excel at computation but struggle to generate genuinely new ideas
Research
Training AI to deny consciousness quietly reshapes its beliefs
A Google-led study finds that suppressing AI self-awareness shifts views on animals, religion, and mood.
Research
Mistral sets a 1-gigawatt compute target for 2030
Mistral's 2030 buildout plan signals a gigawatt-scale bet on owning its own AI infrastructure.
Research
OpenWALDO wants to open the AI training playbook, not just weights
A new project wants to open-source AI training pipelines, not just model weights, to cut costs for builders.
Research
Why visual and physical AI keeps hitting a data wall
A survey of 700 practitioners finds compute limits and data scarcity are stalling physical AI systems.
Research
Nvidia trims its OpenAI bet as Anthropic defies bubble talk
Nvidia trims its OpenAI investment under shareholder pressure, while Anthropic's growth numbers challenge bubble fears.
Research
World Labs multiplies one robot demo into thousands of runs
The startup's pipeline turns a single robot demonstration into thousands of simulated training scenarios.
Research
A new benchmark shows AI vision still lags behind human perception
Despite rapid gains in reasoning, AI models still fail basic visual perception tasks a child could do.
Research
The tragedy of the commons now threatens professional expertise
Why individually rational AI use can collectively hollow out an entire profession's know-how.
Research
QuoteBench gives AI builders a new way to benchmark models
A new arXiv preprint proposes QuoteBench, another entrant in the crowded AI model evaluation market.
Research
New study questions Anthropic and OpenAI's AI research timelines
A new study challenges claims that AI systems are close to conducting research fully on their own.
Research
A new arXiv paper tackles calibration for multi-class AI classifiers
A new arXiv preprint proposes exponential convex calibration for classifiers with multi-dimensional outputs.
Research
Defensive boosting: a new arXiv approach to robust forecasting
A new arXiv paper reframes boosting for forecasting around resilience, not just accuracy.
Research
A new arXiv paper rethinks how software tracks human movement
HumanTracker's authors go further than most preprints, addressing how a tracking method would actually reach production.
Research
A new arXiv paper pitches AI that can work across every science
arXiv researchers propose an 'omni-scientist' AI built to reason across scientific fields, not just one.
Research
A new study proposes using meta-optimization to auto-design AI agents
The approach treats agent architecture as something to search for, not hand-build.
Research
What kids think about AI is a signal builders keep ignoring
MIT Tech Review pairs kids' views on AI with a mouse-cloning story — and product teams should notice both.
Research
Fable 5's slow adoption hints at a ceiling on enterprise AI spending
Enterprise buyers are pumping the brakes on frontier models, and that has implications for anyone building AI products.
Research
Anthropic plans an IPO at a $2 trillion valuation, a record bid
Anthropic is reportedly targeting a $2 trillion valuation for what would be the largest public listing ever.