The Transformer That Wrote Itself a for Loop
When I first read the Coconut paper1 in 2024, I got interested immediately. Their idea is simple: for the intermediate thinking process, you no longer send one token at a time; instead, you feed the model’s own output into the next token’s position, acting as an embedding. Therefore, the model can think in the latent vector space. As a neuroscientist, I’d definitely agree that we think in the latent space, aka the brain, instead of just in words, so how could I not like it? However, the authors acknowledge in the paper that their successful model, though thinking with vectors, was trained with a curriculum that is human-generated, and they tried the ideal case without human supervision, i.e., just questions and answers, but they failed. Since then, this question has kept coming back to me: if a model only sees the final answer, can it develop its own reasoning, in the latent space? I know GRPO and many of its RL descendants do, but I want more of an SFT-like solution. ...