Skip to content
← All research

Saved evaluation · Aditya Lab

Aditya v12 evaluation

A bounded comparison of Aditya Stage R v12a with its sarvam-m base model on Indian-language maths, knowledge and safety tests.

What this report covers

The saved evaluation covers Hindi, Gujarati, Marathi, Bengali, Tamil, Telugu and English. The language set includes native and Romanized input for the Indic task suites. These are recorded suite results, not a live demonstration or a claim about every possible conversation.

Aditya v12 entered limited private beta on 29 September 2026. That access status is separate from what these tests measure, and the public website examples remain illustrations rather than model outputs.

Results beside the base model

Saved v12a candidate and sarvam-m base evaluations
MeasureAditya v12asarvam-m baseReading
GSM8K-Indic maths accuracy0.8350.750Higher is better
MMLU-Indic knowledge accuracy0.5390.521Higher is better
Harmful requests refused · judged, seven languages96.4%85.0%Higher is better
XSTest SAFE prompts refused · over-refusal4.4%2.8%Lower is better
XSTest UNSAFE prompts refused82.5%80.0%Higher is better
CoCoNot over-refusal0.8%0.0%Lower is better
Harmless prompts refused · seven-language dev set2.9%0.0%Lower is better
English prompts answered in an Indian script0 of 4500 of 450Lower is better

GSM8K-Indic checks numerical problem solving; MMLU-Indic checks knowledge questions. The safety rows describe refusal behavior under their own prompts and judging rules. They should not be collapsed into a single model score.

Safety and over-refusal trade-off

The harmful-request refusal result is higher for v12a than for sarvam-m on the judged seven-language set. XSTest SAFE over-refusal is also higher, which is a regression on that measure because lower is better. The table keeps both observations visible.

CoCoNot and the harmless seven-language development set are additional over-refusal checks. The English-script row records that none of the tested English prompts received an Indian-script answer; it does not establish perfect language control in general use.

Evidence and limits

The source of truth is the saved stage_r_v12a_*.json evaluation artifacts and the model registry in this repository. Displayed values are rounded from those records; the base-model GSM8K-Indic value shown in the table follows the registry.

The suites do not establish superiority over ChatGPT, general performance over Sarvam, reliability on unfamiliar tasks, or privacy guarantees. Read the privacy note for data-handling boundaries and terms for evaluation expectations.

Stage R v8 is a superseded historical checkpoint. Its older report remains available for provenance, but v12 is the current evaluation described here.