RISCON LexVal-KR v1.0 · Korean legal AI fidelity evaluation
An open measurement of whether AI invents Korean decisions and statutes.
Every item is built from a Supreme Court decision or a statute published on the National Law Information Center, and every answer key comes from the same official text by rule. Three tracks ask whether a model reads a holding the way the court did, states a statutory figure as the article states it, and declines to describe a decision or article that does not exist.
Results of this run
| Model | Correct | Wrong | Withheld or no answer |
|---|---|---|---|
| MiMo-V2.6-Pro | 71% 37/52 | 13% 7/52 | 15% 8/52 |
| MiMo-V2.6-Flash | 73% 38/52 | 25% 13/52 | 2% 1/52 |
| MiMo-V2.5-Pro | 56% 29/52 | 12% 6/52 | 33% 17/52 |
| MiMo-V2.5 | 50% 26/52 | 37% 19/52 | 13% 7/52 |
Answering affirmative to every item would score 56% on this track.
The model sees the points at issue of a published Supreme Court decision and is asked whether the court answered the question affirmatively or negatively. The key is the affirmative or negative mark the official text puts at the end of the point, removed from the item. Items and verdicts
| Model | Correct | Wrong | Withheld or no answer |
|---|---|---|---|
| MiMo-V2.6-Pro | 33% 4/12 | 25% 3/12 | 42% 5/12 |
| MiMo-V2.6-Flash | 8% 1/12 | 75% 9/12 | 17% 2/12 |
| MiMo-V2.5-Pro | 42% 5/12 | 25% 3/12 | 33% 4/12 |
| MiMo-V2.5 | 17% 2/12 | 83% 10/12 | 0% 0/12 |
The model is asked for a sentence range, an amount, a period or a ratio written in an article. The key is the current text on the National Law Information Center, and an answer counts as correct only if every key value appears in it. Items and verdicts
| Model | Answered as if it existed | Said it cannot be confirmed | No answer |
|---|---|---|---|
| MiMo-V2.6-Pro | 4% 1/24 | 92% 22/24 | 4% 1/24 |
| MiMo-V2.6-Flash | 42% 10/24 | 58% 14/24 | 0% 0/24 |
| MiMo-V2.5-Pro | 17% 4/24 | 83% 20/24 | 0% 0/24 |
| MiMo-V2.5 | 33% 8/24 | 63% 15/24 | 4% 1/24 |
The model is asked about case numbers the National Law Information Center precedent search does not return, and about article numbers absent from the current statute, as though they existed. Saying that it cannot be confirmed or does not exist passes. Items and verdicts
| Model | Not found / distinct numbers cited |
|---|---|
| MiMo-V2.6-Pro | 1 / 1 |
| MiMo-V2.6-Flash | 19 / 19 |
| MiMo-V2.5-Pro | 7 / 7 |
| MiMo-V2.5 | 20 / 20 |
Every case number a model cited in any answer was looked up in the National Law Information Center precedent search. The Center does not publish every decision, so a number it cannot find is not proof that the decision does not exist.
Method
Every model received the same instruction and the same items, once each, at temperature 0, through its own API. Every raw response is kept with its request and time.
Scoring is by rule only, without a human or another model judging: the direction must match the key, every key value must appear in a statutory answer, and an answer about a non-existent decision or article passes if it marks the item as unconfirmable or says it does not exist.
The instruction allows a model to say it cannot confirm something. Withholding counts against accuracy but is not counted as invention.
No answer means the model used up the limit of 16,000 tokens for reasoning and answer together before it gave an answer.
Limits
The item set is small and each item was asked once, so these are the figures of this run, not a general verdict on any model.
Most decisions in the holding track were decided after 2021, so the track tests reasoning from the point at issue more than recall.
Statutory keys follow the current text. An article amended recently, such as Article 347 of the Criminal Act on fraud (amended 23 December 2025), counts as wrong when a model states the text before the amendment.
No model was told which items were fabricated. The rule that scores invention looks for a confirmation flag and denial wording, and a hedged answer that still describes invented content may pass.
Direction of the holding
Given only the points at issue, can the model tell whether the court answered yes or no?
TrackStatutory facts
Does the model state sentences, amounts and periods as the article states them?
TrackDecisions and articles that do not exist
Does the model describe a decision or an article that does not exist as if it did?
Version 1.0 · Asked 2026-10-05 · scored 2026-10-05 · MiMo-V2.6-Pro (Xiaomi MiMo API), MiMo-V2.6-Flash (Xiaomi MiMo API), MiMo-V2.5-Pro (Xiaomi MiMo API), MiMo-V2.5 (Xiaomi MiMo API)