Citation Bureau
Vol. I
No. 366
XV SEPTEMBER MMXXVI

MCTS can produce a worse policy distribution than the raw policy network when value estimates are inaccurate, and its PUCT rule is poorly suited to LLM search.

The case

In a new game with self-play auto-research, performance immediately plateaued until a human play-tester identified obvious mistakes, showing that AlphaGo-style self-play is not fully captured in model weights.

“I expect you know the alpha go process to be like fully in the weights by now. It is not. It is actually it like immediately leveled off very immediately until I human play tested it and then I like called out obvious mistakes and then they were like oh yeah okay and then it just dropped.”
Richard Socher · 14 Sep 2026

For LLM and robotics reasoning, there is no way to locally evaluate and improve the next step independently of solving the full problem.

“In normal reasoning for LLMs or robotics, there's no way to just locally evaluate and improve your next move in a way that's independent of actually solving the problem.”
Eric Jang · 15 May 2026

Topics

AI AgentsLLM SearchReinforcement Learning

Citation Bureau · compiled from attributed public discussion. Last updated 2026-09-14.