New method strengthens LLM reasoning without human supervision
Researchers introduce FREIA, a new algorithm that enhances the unsupervised reasoning capabilities of large language models (LLMs) through adaptive reinforcement. The method particularly excels in mathematical tasks.

What happened?
A new algorithm called FREIA (Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping) has been presented, aimed at improving unsupervised reinforcement learning (RL) for large language models (LLMs). FREIA addresses deficiencies in existing methods through two key innovations: Free Energy-Driven Reward (FER), which balances consensus and exploration in rewards, and Adaptive Advantage Shaping (AAS), which adjusts learning signals based on the statistical properties of the rewards.
Key facts
”Unsupervised reinforcement learning (RL) has emerged as a promising paradigm for enabling self-improvement in large language models (LLMs). However, existing unsupervised RL-based methods often lack the capacity to adapt to the model's evolving reasoning capabilities during train”
”To address this issue, we introduce FREIA, a novel RL-based algorithm built on two key innovations: (1) Free Energy-Driven Reward (FER) adapts rewards to balance consensus and exploration based on the Free Energy Principle. (2) Adaptive Advantage Shaping (AAS) adaptively adjusts”
”Empirical evaluations on nine datasets across three reasoning tasks showcase that FREIA outperforms other unsupervised RL-based baselines. Notably, in mathematical reasoning tasks, FREIA surpasses other me”
Why it matters
Current unsupervised RL methods for LLMs often lack the ability to adapt to the model's evolving reasoning capabilities during training. This can lead to suboptimal policy optimisation in the absence of ground-truth data. FREIA's adaptive approach aims to bridge this gap, enabling more efficient and autonomous improvement of LLM reasoning capacity.
Who is affected?
LLM developers and machine learning researchers are the primary beneficiaries of this research, as FREIA provides a tool to streamline and enhance the training of language models. Users benefit indirectly from more capable and reliable AI systems, particularly in areas requiring complex reasoning.
What else you should know
Empirical evaluations on nine datasets across three reasoning tasks show that FREIA outperforms other unsupervised RL-based baselines. The algorithm demonstrates particularly strong results in mathematical reasoning tasks.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vilka bolag berörs?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "New method strengthens LLM reasoning without human supervisi"