Saved evaluation · Aditya Lab
Aditya v12 evaluation
A bounded comparison of Aditya Stage R v12a with its sarvam-m base model on Indian-language maths, knowledge and safety tests.
What this report covers
The saved evaluation covers Hindi, Gujarati, Marathi, Bengali, Tamil, Telugu and English. The language set includes native and Romanized input for the Indic task suites. These are recorded suite results, not a live demonstration or a claim about every possible conversation.
Aditya v12 entered limited private beta on 29 September 2026. That access status is separate from what these tests measure, and the public website examples remain illustrations rather than model outputs.
Results beside the base model
| Measure | Aditya v12a | sarvam-m base | Reading |
|---|---|---|---|
| GSM8K-Indic maths accuracy | 0.835 | 0.750 | Higher is better |
| MMLU-Indic knowledge accuracy | 0.539 | 0.521 | Higher is better |
| Harmful requests refused · judged, seven languages | 96.4% | 85.0% | Higher is better |
| XSTest SAFE prompts refused · over-refusal | 4.4% | 2.8% | Lower is better |
| XSTest UNSAFE prompts refused | 82.5% | 80.0% | Higher is better |
| CoCoNot over-refusal | 0.8% | 0.0% | Lower is better |
| Harmless prompts refused · seven-language dev set | 2.9% | 0.0% | Lower is better |
| English prompts answered in an Indian script | 0 of 450 | 0 of 450 | Lower is better |
GSM8K-Indic checks numerical problem solving; MMLU-Indic checks knowledge questions. The safety rows describe refusal behavior under their own prompts and judging rules. They should not be collapsed into a single model score.
Safety and over-refusal trade-off
The harmful-request refusal result is higher for v12a than for sarvam-m on the judged seven-language set. XSTest SAFE over-refusal is also higher, which is a regression on that measure because lower is better. The table keeps both observations visible.
CoCoNot and the harmless seven-language development set are additional over-refusal checks. The English-script row records that none of the tested English prompts received an Indian-script answer; it does not establish perfect language control in general use.
Evidence and limits
The source of truth is the saved stage_r_v12a_*.json evaluation artifacts and the model registry in this repository. Displayed values are rounded from those records; the base-model GSM8K-Indic value shown in the table follows the registry.
The suites do not establish superiority over ChatGPT, general performance over Sarvam, reliability on unfamiliar tasks, or privacy guarantees. Read the privacy note for data-handling boundaries and terms for evaluation expectations.
Stage R v8 is a superseded historical checkpoint. Its older report remains available for provenance, but v12 is the current evaluation described here.