Princeton's Mengdi Wang Tells Bund Conference AI Needs Verifiable Science, Not Just Bigger Models
Princeton's Mengdi Wang told the Bund Conference that large models return the most likely answer rather than the long tail where discoveries lie, and said verifiable experiments matter more than scale.
Wang's argument turns on where discoveries sit in a probability distribution: models are good at returning the most likely answer, while genuinely new findings tend to sit in the long tail. Her team, working with Princeton sociology professor Yu Xie, tested this by asking ChatGPT and Claude to simulate a life. In one experiment the model was asked to play a British person born in 1930 and generate that person's subsequent trajectory on its own. Across questions about marriage age, occupation and gender, the models overestimated the most common outcomes and clearly underestimated minority cases and the long tail.
"The model learns the maximum likelihood point of a probability distribution, but it may underestimate or even ignore long-tail knowledge," Wang said.
She used this to explain why models have advanced quickly in mathematics and coding but have not produced comparable original discoveries in physics, chemistry or biology. Mathematics and coding have relatively clear verifiers: compilers and unit tests for code, formal proof systems such as Lean for mathematics. A model can try, receive feedback, correct itself and enter the next training round. Verification in real science is far harder.
"Ideas are cheap. You have to show me the code, or show me the experiment," Wang said. Scientific research involves complex equipment, human judgment and cross-team collaboration, and many experiments cannot be reproduced reliably. More information does not mean more effective new information, she said, and can instead lower the signal-to-noise ratio of the research process.
In her view, the key to AI-driven autonomous discovery is not giving models more knowledge but making the real world increasingly verifiable. Her team is building infrastructure for that. In quantum materials research it has set up an automated experimental platform that runs graphene experiments which previously took researchers months, and packages experimental equipment and capabilities as APIs so other research institutions can call and reproduce them. The team is also exploring LabOS, an intelligent operating system for research laboratories that would let multimodal AI perceive the experimental environment through smart glasses, assist researchers in their operations and connect software, AI models, robots and real experimental equipment.
Wang said what determines whether AI becomes a discoverer is not simply model size, but whether infrastructure exists that links hypothesis, verification, traceability and repetition. Only when physical experiments can be systematically recorded, continuously verified and repeatedly reproduced can AI turn the likelihood it learned in training into exploration of what is possible.
The 2026 Inclusion·Bund Conference runs from Sept. 9 to 12 in Shanghai under the theme "Co-creating the AI New Economy," discussing how AI can drive more sustainable economic growth and broader social value.