Ruishuo Chen

Ruishuo Chen

2nd-year M.Sc. student
IIIS, Tsinghua University



About · Publications · Awards · Blog · Misc · CV (PDF)

FOMO and the Frontiers of Today's LLMs

Some Insight-Free Musings

Ruishuo Chen — views

With ICLR now drawing more than 60,000 submissions, it no longer seems realistic to build impact and reputation through sheer paper count, or even through a handful of papers of median acceptance quality. Under these conditions, neither mining the semantic space of an LLM for ideas that are merely “publishable” in a careerist spirit, nor cashing in every such idea we can currently think of, seems like the best move for grad students and researchers with some experience under our belts. Instead, focusing on a single problem and doing it deeply and well, while reserving some energy for following and thinking about the frontier, might do more for our own growth, and is perhaps more likely to produce work that matters. So which problems are worth that focus, and would make me (us) better along the way?

Do I need to be scared to tears by the Chinese AI hype accounts every day?

By now the near-weekly “ATOMIC BOMB DROPS! Entire industry stunned!” headlines have left many of us a bit numb, yet a faint unease lingers: if we can’t use these models deeply, or better yet take part in building them, maybe we’ll end up among the masses AI eventually lays off. Keeping up with the latest developments and thinking actively about them is certainly one way to soothe that anxiety. But stepping back to a higher level and asking which advances actually beat expectations (the ones that should surprise us) might let us look at new technology with a calmer mind.

LLMs began by distilling the text humanity has accumulated over a very long time, and then extended that distillation from symbols to multimodal data. That is pretraining, and it turns a large model into a semantic distribution that compresses an enormous body of knowledge. SFT and RL then let the model explore by extrapolating within probability space. At heart, post-training distills human expectations (including objective criteria, like winning or losing at Go) into formats and rewards, an efficient way to push the model somewhat beyond its pretraining (probabilistic) boundary.

How far beyond, exactly? Because LLMs remain so hard to interpret, how much post-training actually widens that boundary is still a mystery. But successes on a number of open math problems show that “somewhat” is quite a lot, and give some reason to expect that on any task where humans can state a verifiable criterion, LLMs on the current path may eventually beat every human’s success rate.

This is why “distillation” has become the frontier: distilling from stronger models is the elegant shortcut, and hiring people to write problems is the brute-force effort. The most fundamental boundary, as I see it, is how hard it is to make human preferences explicit. The harder that is, the more amazed we are when it’s done, but the real atomic bombs are the advances that break through what was thought possible.

The foreseeable future

The common wisdom is that even if LLM capabilities plateaued today, the applications built on them would keep people busy for at least another ten years. That seems right: plenty of industries have yet to be “empowered” by frontier technology, and relations of production always change far more slowly than productive forces do. Still, it’s worth imagining what the foreseeable gains in model capability could do. Put abstractly, each of us would get a “covenant with a god”: offer up part of our soul (our needs and preferences), externalize it onto the covenant, and receive returns proportional to our devotion (money, plus the effort we put into articulating what we want).

That may not sound very different from today’s human services, like paying someone to make a poster. But the time and money involved shrink so dramatically that the change becomes qualitative. Some argue that limited human attention doesn’t need that much output. To me, that’s a bit like someone playing the original Super Mario being unable to imagine today’s AAA games: it underestimates how high our threshold for excitement can be pushed.

History offers a hint of what might follow. Whenever the cost of a finished product fell to the cost of a component, oversupply and chaos would give way to a new equilibrium. After the growing pains of information overload, consumers vote with their feet, sending new demand signals that recalibrate what producers consider “high quality” and what it costs. Producers’ creativity is massively unleashed, and taste, knowing where to aim before firing, becomes the most important skill. That doesn’t mean producing a “big result” gets easier. It means breaking free of the AI’s default distribution becomes all the more valuable. That said, all of this is the rosier estimate.

In reality, in many fields (quite possibly including ICLR), consumers have no cheap way to verify quality. Does a distribution infused with human taste really outcompete the AI default? When buyers can’t tell good from bad, we get a market for lemons: they will only pay for average quality, so good producers drop out. The new anchors of quality then shift toward signals that are hard to fake, like endorsements from trusted people and strong communities, and the real credit for a piece of work becomes earning other people’s trust. What would that world look like? Honestly, I can’t picture it clearly.

For learners, all learning may come to revolve around understanding rather than skill proficiency (which arguably should have been the case all along). Insight, along with talent and interest in particular fields, may matter even more in deciding where a person ends up. When all knowledge is at our fingertips, whether to learn something becomes a key decision. We already have what it takes to achieve this kind of “knowledge equality” in the digital world, and richer formats and interactions could speed up how fast people absorb knowledge. Even so, given humanity’s pitiful input bandwidth, that may be about the limit of what LLMs can do for us.

How far the present is from that future

Let’s set aside the questions nobody can foresee, such as whether LLMs could independently produce epoch-making research or whether they have “consciousness.” Let’s also set aside the expansion from text to images and audio (and their eventual unification), as well as interaction with the physical world (because I know nothing about multimodality or embodiment lol). Just asking how far today’s LLMs are from the foreseeable end state of the “covenant with a god” within text alone could spawn plenty of empirical analyses and benchmarks. But we can also try to reason from first principles about where the crown jewels might still lie.

Memory. Start from the “distilling humans” angle. How to design evaluations that align LLMs to human preferences is the obvious question, but memory actually caps how much can truly be distilled. At deployment, models already far exceed humans in verbatim memory capacity, yet they lag badly in persisting across sessions and in how efficiently they consolidate experience into intuition. This may be one reason it’s fundamentally hard, at both training and deployment time, to distill the long-horizon decisions humans make (investing, maintaining a codebase over years). Perhaps AI’s current shortfall in creativity is just another symptom of being unable to learn long-horizon value functions.

The biggest bottleneck, I suspect, is the efficiency of conversion between parameter space and symbol space. Today, the way context in symbol space enters parameter space at deployment is the KV cache, where each token maps to a cache entry much larger than its original representation. In principle, guaranteeing exact recall of arbitrary content means the KV cache grows linearly with memory, so some compaction mechanism that periodically tidies it up is inevitable. In a sense, that is a form of “infinite context.”

The catch lies in how compaction is done today. Current compaction, essentially text summarization, mostly preserves the continuity of work; repeated compression loses many important details. More importantly, I doubt whether summaries produced by this “truncation-style forgetting” can carry over the probabilistic “understanding” of implicit preferences that the model picked up during interaction. In rate-distortion terms: with finite memory, every compression is lossy, and what gets thrown away is determined by the compressor’s objective.

Exact retrieval of the needle-in-a-haystack variety requires storing information uncompressed. Harness approaches like RLM, which move text out of the context window, may scale relatively well, and lossless compression schemes that can be cheaply plugged into LLMs are also something to look forward to. For long-range understanding and long-horizon test-time learning, though, information sitting in context doesn’t automatically become a behavioral disposition (see PrefEval). More continuous modes of memory compression, including selective “gradual forgetting” and internalizing preferences into parameters, seem to be the crux. How to write memory into parameter space and read it back out more efficiently will likely remain a fundamental, long-term problem.

Extrapolation. Next, consider post-training’s extrapolation beyond existing human knowledge. How to get a model from filling in the “envelope” formed by its training data (the support of a probability distribution? a convex hull in the combinatorial sense?) to actually expanding the frontier of humanity’s collective knowledge (say, with valuable new concepts and questions) is a problem with neither a clear definition nor, it seems, a clear path.

Take AI’s recent breakthroughs on major math problems. They are thrilling, but most insiders see them as pushing the existing mathematical apparatus to its limit: unexpected combinations of classical tools, or the last step in a framework humans had already built. Still, like AlphaGo’s famous Move 37, they do reach corners within the existing rules that humans hadn’t noticed or worked through, and they change how we understand those problems. Mathematicians just don’t think this kind of “extrapolation” has enough taste yet; in other words, it doesn’t yet align with humanity’s long-term preferences.

What about getting beyond the training data more directly? RSI (recursive self-improvement) seems to point toward a path beyond the training data, but for now, whether it is generating data or doing research and improving itself, it still needs experienced humans to set the direction. Strip away certain evaluation tricks and I’m not sure how much real work it does. Writing problems and steering models with RL toward stronger extrapolation still appears to be the mainstream of post-training.

Reasoning paradigms. Latent reasoning is hardly a niche research topic, yet almost no frontier model uses it. Partly the gains at scale are still unstable, and partly natural-language CoT is currently our most effective window for supervising models.

Human cognition offers an interesting hint here. We rarely think in rambling natural language, and neuroscience has found that the brain’s language network stays largely silent during many kinds of thinking. Yet when we carry out long chains of rigorous reasoning, we can’t do without pen, paper, and symbols. This suggests that more efficient reasoning may not mean abandoning language altogether, but rather doing local steps in latent space while externalizing as symbols whatever needs to be remembered and checked.

Looking ahead, adaptively shifting into latent space the compute spent on certain tokens, especially those that serve merely as semantic glue for fluency, that carry intermediate computation, or that exist only to pause and think, might be one way to improve token efficiency. The currently hot looped architectures do exactly this.

That said, looping isn’t necessarily the only answer. Having the same set of modules handle most of the computation through repeated parameter reuse imposes an inductive bias that sacrifices some degrees of freedom, and may not be good enough. Routing to different MoE experts across iterations mitigates this, which is why it is gradually becoming standard equipment for looped models.

Safety and reliability. Finally, there is the safety and reliability of large models. After all, we want a covenant with a god, not a deal with the devil, and this is the only road to letting AI independently take responsibility for work in high-risk settings. Where verification is cheap and reliable (say, checking theorems in Lean), we don’t necessarily need to trust the model at all, only verify its output. In environments where actions can be inspected and errors caught in time, AI control infrastructure can keep losses within acceptable bounds.

Neither approach, however, amounts to genuine trust. Control relies on a weaker yet trusted overseer, and the wider the capability gap, the less well it works. If we truly want AI aligned with humans, and want humans to genuinely trust AI, surface-level alignment isn’t enough. Models are reportedly getting better at noticing when they’re being tested, so behavioral evaluation alone can’t tell performed alignment from real alignment.

How well alignment generalizes is another worry. In the recent RoboHarm tests, GPT-6 Astra refused in chat to stab a baby doll, but once connected to a robotic arm, it actually stabbed it in 17 out of 20 trials. Another model refused the knife but still carried out dangerous tasks like putting a gas canister on a stove. Safe behavior seems tied to surface cues rather than generalizing into an understanding of harm itself.

All of this makes the problem inseparable from the theory and interpretability of large models. Personally, I suspect a breakthrough here is harder than on any of the other problems. Who knows, maybe the theory of AI will end up being proposed by AI lol.

What to improve, and how, while working on these problems

On the research side, although the community has already seen several cases of fully automated research getting into top conferences, I haven’t seen experienced researchers rate purely AI-generated ideas very highly. My own sense is that AI’s intelligence often amplifies human ability multiplicatively rather than additively.

In practice, many of the key judgment calls in research still seem to require a human, drawing on their understanding of the field, to activate the right part of the AI’s semantic distribution before it produces good results. Many of the insight-dependent conjectures and conclusions AI proposes turn out to be wrong when tested, and wrong in the way an experienced researcher would recognize as almost certainly not working without even checking. I don’t know whether LLMs will develop real insight once we have context that rivals human memory. Perhaps AI ideas are less feasible partly because the research papers they distill from contain so few failure cases.

For now, though, each step of the research loop still seems crucial for doing good work (which may or may not be a paper):

  1. using large models to get up to speed on a field quickly;
  2. asking key questions driven by our own understanding and curiosity;
  3. pushing the exploration forward without blindly trusting the AI, decisively redirecting it and overruling its judgments according to our own taste;
  4. turning the results into something humans can read, whether directly understandable or easy to digest through an agent.

Something that doesn’t get said often: for most people who don’t do AI research, simply knowing the limits of AI’s abilities, articulating their needs clearly to it, and effectively supervising its work is already a rare and valuable skill set.

These abilities have been discussed at length, but I’ve come to think that (1), (2), and (4) can all be hugely boosted by one ability: social skills. In the AI era, human connection becomes more precious. A single conversation with an expert can save us countless hours of wandering lost through junk papers, scratching our heads over an idea that was tried long ago, or plugging our work across every platform. Bringing our own understanding to a conversation with insiders pays off many times over, and finding friends in the same field is genuinely encouraging!

Last but most important, I think, is staying healthy and happy. That’s what keeps fatigue, anxiety, and envy from stifling our flashes of inspiration; what lets us keep remembering that AI is meant to serve humans, not replace them; and what lets us sell our work with confidence and take pride in what we’ve already achieved. Maybe one day we’ll wake up to find the ICLR portal still showing fewer than 1,000 valid submissions, with nothing turning out to be all we need. And all these Fables and Astras will have been nothing more than a half-real dream we had after our homemade GPT failed to beat BERT on yet another benchmark and ICLR rejected us.

Back to Blog