New method reduces costs for conversational AI using small models
Researchers present a new method to streamline multi-turn AI dialogues by using smaller models to reduce costs and latency without compromising quality.

What happened?
A recently published paper on arXiv describes a new method aimed at making the operation of Large Language Models (LLMs) in dialogue systems more efficient. The method is based on initially identifying a "local response manifold" during the first dialogue steps. Subsequently, a smaller surrogate model is adapted to this specific context to handle the remainder of the conversation.
Key facts
| Publikationsdatum | 23 maj 2024 |
|---|---|
| Typ av modell | LLM (Large Language Model) och Small Language Model |
| Källpublikation | arXiv cs.CL |
”A standard serving practice concatenates the full dialogue history at every turn, which reliably maintains coherence but incurs substantial cost in latency, memory, and API expenditure, especially when queries are routed to large proprietary models.”
”We propose a framework that exploits the early turns of a session to estimate a local response manifold and then adapt a smaller surrogate model to this local region for the remainder of the conversation.”
”Concretely, we learn soft prompts that maximize semantic divergence between the large and surrogate small language models' responses to surface least-aligned local directions, stabilize training with anti-degeneration control, and distill th”
Why it matters
Current standard practice of sending the entire conversation history to large models at every step is inefficient in terms of both cost and computation. By distilling knowledge from the large model into a smaller, specialised model for a specific dialogue context, significant savings in latency and memory usage can be achieved. This is particularly relevant for AI assistant services and chatbots where many users interact with the systems simultaneously.
Who is affected?
The method primarily affects developers and companies hosting conversational AI systems using LLMs. Users may indirectly benefit from faster response times and potentially lower costs for AI services. Researchers in LLM optimisation and machine learning are also a direct target audience for this type of research.
What else you should know
The method utilises "soft prompts" to maximise the semantic divergence between the large model and surrogate model responses. This identifies the local areas where the surrogate model requires extra training, and stabilises training with "anti-degeneration control" to prevent the model's performance from deteriorating over time.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vilka bolag berörs?
Påverkar det EU?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "New method reduces costs for conversational AI using small m"