FormalASR converts spoken Chinese to formal text in real time
Researchers have developed FormalASR, an end-to-end AI model that converts spoken Chinese directly into formal written text, reducing the need for post-editing.

What happened?
A new AI model, FormalASR, has been launched. It is designed to directly transcribe spoken Chinese into formal written text. Unlike traditional ASR systems, FormalASR avoids including colloquial elements such as hesitations and informal phrases, which often require extensive post-editing. The models, available in two sizes (0.6B and 1.7B parameters), are trained on two new datasets: WenetSpeech-Formal and Speechio-Formal.
Key facts
| Modellnamn | FormalASR |
|---|---|
| Modellstorlekar | 0.6B och 1.7B parametrar |
| Förbättring (relativ CER-minskning) | Upp till 37.4% |
| Datum för publicering | 24 maj 2026 |
”Automatic speech recognition (ASR) systems are typically optimized for verbatim transcription, which preserves disfluencies, filler words, and informal spoken structures that are often unsuitable for downstream writing-oriented applications.”
”We present FormalASR, two compact end-to-end models (0.6B and 1.7B) that directly transcribe spoken Chinese into formal written text.”
Why it matters
Traditional Automatic Speech Recognition (ASR) systems focus on verbatim transcription, which often results in text unsuitable for writing-oriented applications. The standard two-step solution (ASR followed by an LLM for post-editing) introduces latency and increased memory consumption. FormalASR addresses this by providing direct conversion, streamlining the process and enabling easier deployment on devices.
Who is affected?
Developers of AI and language models can benefit from the end-to-end architecture offered by FormalASR. Companies working in transcription, customer service, and medical dictation may see improved efficiency and reduced post-processing costs. Users requiring fast and accurate captioning of spoken content are positively impacted.
What else you should know
FormalASR demonstrates a relative reduction in Character Error Rate (CER) of up to 37.4% compared to existing verbatim ASR systems.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vilka språk stöder FormalASR?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "FormalASR converts spoken Chinese to formal text in real tim"