New method teaches AI models self-criticism via reinforcement learning
Researchers introduce ICRL, a new framework that enables large language models to internalise self-criticism to improve performance without external guidance.

What happened?
Researchers have published a new method called ICRL (Internalise Self-Critique with Reinforcement Learning). This method trains a solver and a critic jointly from a shared base to transform successes derived from criticism into a capability the solver can use independently. Critics are rewarded based on the solver's subsequent performance improvement, encouraging actionable feedback.
Key facts
| Metod | ICRL (Internalize Self-Critique with Reinforcement Learning) |
|---|---|
| Typ av AI | Stora språkmodeller (LLMs), Agentbaserade AI-system |
| Publicerad | 2605.15224v1 (förmodligen 15 maj 2026) |
”Large language model-based agents make mistakes, yet critique can often guide the same model toward correct behavior. However, when critique is removed, the model may fail again on the same query, indicating that it has not internalized the critique's guidance into its underlying”
”To address this, we propose learning to internalize self-critique with reinforcement learning(ICRL), a novel framework that jointly trains a solver and a critic from a shared backbone to convert critique-induced success into unassisted solver ability.”
Why it matters
Large language models often make mistakes but can be corrected with external criticism. The problem is that models often fail again when criticism is removed, suggesting they have not internalised the guidance. ICRL aims to solve this by allowing the model to learn from its own criticism, potentially leading to more robust and reliable AI systems.
Who is affected?
This primarily affects AI researchers and developers working with large-scale language models and agent-based AI systems. In the long run, it could benefit all users of AI applications by leading to more reliable and efficient language models.
What else you should know
ICRL also introduces a reweighting ratio for distribution calibration that selectively chooses the most helpful criticism steps via reinforcement learning to manage the difference in behaviour between conditional and unconditional criticism.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vem påverkas främst av detta?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "New method teaches AI models self-criticism via reinforcemen"