JOR-Bench evaluates LLM performance in Japanese optimisation
A new set of five Japanese benchmarks, JOR-Bench, has been introduced to evaluate large language models' ability to formulate and solve optimisation problems.

What happened?
JOR-Bench is a collection of five Japanese-language benchmarks for evaluating the performance of large language models (LLMs) in formulating and solving Operations Research (OR) problems. The benchmark consists of 1,319 problems translated from existing English benchmarks such as IndustryOR, MAMO Complex LP, NL4OPT, OptiBench, and OptMATH. The problems cover linear programming, mixed-integer optimisation, non-linear programming, and combinatorial optimisation.
Key facts
| Antal problem | 1 319 |
|---|---|
| Antal benchmarks inkluderade | 5 |
| Datum för publicering | 25 juli 2026 |
| Områden optimering | Linjär, blandad heltal, icke-linjär, kombinatorisk |
| Antal utvärderade LLM:er | 7 |
”We present JOR-Bench, a collection of five Japanese-language benchmarks for evaluating the ability of large language models (LLMs) to formulate and solve operations research (OR) problems.”
Why it matters
JOR-Bench provides a standardised method for evaluating LLM performance in Operations Research within a Japanese context. It enables comparisons between various models, including multilingual and Japanese-specialised models, in both English and Japanese. This facilitates the development of more capable LLMs that can handle complex optimisation problems in practical applications.
Who is affected?
Researchers, large language model developers, and companies using LLMs for optimisation problems are affected. Specifically, Japanese researchers and companies can benefit from the new benchmarks to assess and improve their models' performance.
What else you should know
JOR-Bench is solver-agnostic and can be used with any optimisation solver or programming language. The evaluation in the study was standardised using the Python interface for OR-Tools.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vilka typer av problem täcker JOR-Bench?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Which processes can be simplified or automated based on this?
- Who trains the team — and when? Set a clear owner and deadline.
- Follow up KPIs on lead time, quality and cost after adoption.
Generated angle — not editorial analysis of "JOR-Bench evaluates LLM performance in Japanese optimisation"