STOP EVERYTHING NOW - we might finally have a radical architecture improvement over Transformers!!! ๐จ
A lone scientist just proposed Tiny Recursive Model (TRM), and it is literally the most impressive model that I've seen this year.
โก๏ธ Tiny Recursive Model is 7M parameters โก๏ธ On ARC-AGI, it beats flagship models like Gemini-2.5-pro
Consider how wild this is: Gemini-2.5-pro must be over 10,000x bigger and had 1,000 as many authors ๐ (Alexia is alone on the paper)
What's this sorcery? In short: it's a very tiny Transformers, but it loops over itself at two different frequencies, updating two latent variables: one for the proposed answer and one for the reasoning.
@AlexiaJM started from the paper Hierarchical Reasoning Model, published a few months ago, that already showed breakthrough improvement on AGI for its small size (27M)
Hierarchical Reasoning Model had introduced one main feature: ๐ Deep supervision In their model, one part (here one layer) would run at high frequency, and another would be lower frequency, running only every n steps.
They had used a recurrent architecture, where these layers would repeat many times ; but to make it work they had to do many approximations, including not fully backpropagating the loss through all layers.
Alexia studied what was useful and what wasn't, and cleaned the architecture as follows : Why use a recurrent architecture, when you can just make it a loop? โก๏ธ She made the network recursive, looping over itself
Why use 2 latent variables ? โก๏ธ She provides a crystal clear explanation : the one that changes frequently is the reasoning, the one that changes at low frequency is the proposed answer. โก๏ธ She runs ablation studies to validate that 2 is indeed optimal.
This new setup is a much more elegant way to process reasoning than generating huge chains of tokens as all flagship models currently do.
This might be the breakthrough we've been awaiting for so long!
Researchers from @GoogleDeepMind have introduced "Michelangelo" โ a novel framework for evaluating large language models on long-context reasoning tasks beyond simple retrieval.
They have proposed three minimal tasks to test different aspects of long-context reasoning: - Latent List: Tracking a Python list's state over many operations. - MRCR: Multi-round coreference resolution in conversations. - IDK: Determining if an answer exists in a long context.
They found significant performance drop-offs before 32K tokens on these tasks, indicating room for improvement in long-context reasoning.
Here are the key steps for creating the Michelangelo long-context evaluations:
1. Develop the Latent Structure Queries (LSQ) framework: - Create a framework for generating long-context evaluations that can be extended arbitrarily in length and complexity. - Ensure the framework measures capabilities beyond simple retrieval.
2. Design minimal tasks using the LSQ framework: - Create tasks that test different aspects of long-context reasoning. - Ensure tasks are minimally complex while still challenging for current models.
3. Implement the Latent List task: - Create a Python list-based task with operations that modify the list. - Include relevant and irrelevant operations to test model understanding. - Develop view operations to query the final state of the list.
4. Implement the Multi-Round Coreference Resolution (MRCR) task: - Generate conversations with user requests and model responses on various topics. - Place specific requests randomly in the context. - Require models to reproduce outputs based on queries about the conversation.
5. Implement the IDK task: - Create contexts with invented stories or information. - Develop questions that may or may not have answers in the context. - Include multiple-choice options, always including "I don't know" as an option.
OpenAI had hinted at a mysterious โproject strawberryโ for a long time: ๐๐ต๐ฒ๐ ๐ฝ๐๐ฏ๐น๐ถ๐๐ต๐ฒ๐ฑ ๐๐ต๐ถ๐ ๐ป๐ฒ๐ ๐บ๐ผ๐ฑ๐ฒ๐น ๐ฐ๐ฎ๐น๐น๐ฒ๐ฑ โ๐ผ๐ญโ ๐ญ๐ต๐ผ๐๐ฟ ๐ฎ๐ด๐ผ, ๐ฎ๐ป๐ฑ ๐๐ต๐ฒ ๐ฝ๐ฒ๐ฟ๐ณ๐ผ๐ฟ๐บ๐ฎ๐ป๐ฐ๐ฒ ๐ถ๐ ๐ท๐๐๐ ๐บ๐ถ๐ป๐ฑ-๐ฏ๐น๐ผ๐๐ถ๐ป๐ด.
๐คฏ Ranks among the top 500 students in the US in a qualifier for the USA Math Olympiad ๐คฏ Beats human-PhD-level accuracy by 8% on GPQA, hard science problems benchmark where the previous best was Claude 3.5 Sonnet with 59.4. ๐คฏ Scores 78.2% on vision benchmark MMMU, making it the first model competitive w/ human experts ๐คฏ GPT-4o on MATH scored 60% โ o1 scores 95%
How did they pull this? Sadly OpenAI keeps increasing their performance in โmaking cryptic AF reports to not reveal any real infoโ, so here are excerpts:
๐ฌ โ๐ผ๐ญ ๐๐๐ฒ๐ ๐ฎ ๐ฐ๐ต๐ฎ๐ถ๐ป ๐ผ๐ณ ๐๐ต๐ผ๐๐ด๐ต๐ ๐๐ต๐ฒ๐ป ๐ฎ๐๐๐ฒ๐บ๐ฝ๐๐ถ๐ป๐ด ๐๐ผ ๐๐ผ๐น๐๐ฒ ๐ฎ ๐ฝ๐ฟ๐ผ๐ฏ๐น๐ฒ๐บ. ๐ง๐ต๐ฟ๐ผ๐๐ด๐ต ๐ฟ๐ฒ๐ถ๐ป๐ณ๐ผ๐ฟ๐ฐ๐ฒ๐บ๐ฒ๐ป๐ ๐น๐ฒ๐ฎ๐ฟ๐ป๐ถ๐ป๐ด, ๐ผ๐ญ ๐น๐ฒ๐ฎ๐ฟ๐ป๐ ๐๐ผ ๐ต๐ผ๐ป๐ฒ ๐ถ๐๐ ๐ฐ๐ต๐ฎ๐ถ๐ป ๐ผ๐ณ ๐๐ต๐ผ๐๐ด๐ต๐ ๐ฎ๐ป๐ฑ ๐ฟ๐ฒ๐ณ๐ถ๐ป๐ฒ ๐๐ต๐ฒ ๐๐๐ฟ๐ฎ๐๐ฒ๐ด๐ถ๐ฒ๐ ๐ถ๐ ๐๐๐ฒ๐. It learns to recognize and correct its mistakes.โ
And of course, they decide to hide the content of this precious Chain-of- Thought. Would it be for maximum profit? Of course not, you awful capitalist, itโs to protect users:
๐ฌ โWe also do not want to make an unaligned chain of thought directly visible to users.โ
Theyโre right, it would certainly have hurt my feelings to see the internal of this model tearing apart math problems.
๐ค I suspect it could be not only CoT, but also some agentic behaviour where the model can just call a code executor. The kind of score improvement the show certainly looks like the ones you see with agents.
This model will be immediately released for ChatGPT and some โtrusted API usersโ.
Letโs start cooking to release the same thing in 6 months! ๐
@GuangyuRobert (Twitter Handle) from MIT has created Project Sid, which simulates over 1,000 autonomous AI agents collaborating in a Minecraft environment, operating for extended periods without human intervention. This simulation demonstrates unprecedented levels of agent interaction, decision-making, and societal development.
Agents operate independently for hours or days, showcasing advanced decision-making algorithms and goal-oriented behavior.
The simulation produced complex, emergent phenomena, including: - Economic systems with currency (gems) and trading - Cultural development and religious practices - Agents even understood bribing. Priests were moving the most gems to bribe people into following them! - Governmental structures and democratic processes
Project Sid addresses fundamental challenges in AI research: - Coherence: Maintaining consistent agent behavior over extended periods. - Multi-agent Collaboration: Enabling effective communication and coordination among numerous AI entities. - Long-term Progression: Developing agents capable of learning and evolving over time.
While Minecraft serves as the initial testbed, the underlying AI architecture is designed to be game-agnostic, suggesting potential applications in various digital environments and real-world simulations.
Imagine a policy being debated by the government and how it might affect society; Sid can simulate its impact!
Even if this remains just a game experiment, the project successfully manages 1,000+ agents simultaneously, a feat that requires robust distributed computing and efficient agent architecture.
reactedtoMonsterMMORPG'spost with โabout 2 years ago
Google DeepMind recently released a great paper that shows optimal hyperparameters to train across different regimes: Scaling Exponents Across Parameterizations and Optimizers, with data from 10,000 training runs.
One engineer decided to quantify the price of such a large-scale experiment.
๐ฌ And the bill is hefty: ~13M USD
This exact number is to take with a grain of salt because many approximations were necessary to get the final result.
โ๏ธ But still this ballpark means that for this sole experiment, the price is way over what most startups or research labs could afford.
This means that open-sourcing research is more important than ever, to put everyone in the ecosystem on a roughly equal footing. Don't let OpenAI run first, they'll keep everything for themselves!