Stage R v8: reasoning gains across native and Romanized Indic
Recorded GSM8K-Indic and MMLU-Indic results, the language slices, and the regression that a single score would hide.
Historical results from a deprecated checkpoint, not a new inference run. This is an evaluation report, not a peer-reviewed paper or release certification.
The short version
On the saved 1,300-row GSM8K-Indic evaluation, Stage R v8 answered 1092 items correctly, compared with 963 for the corrected Sarvam-M base run. That is 84.00% versus 74.08%, a gain of 9.92 percentage points on this suite.
The MMLU-Indic change was smaller: 52.12% to 53.96% on 2,600 multiple-choice rows. Native-script Telugu moved in the wrong direction. The evidence supports a task-specific maths improvement, not a claim that the fine-tune improved every language or every kind of reasoning.
These figures come from repository artifacts, not the website’s interface illustrations. We recalculated the aggregate and per-language counts from the saved row-level results before building the figures. We did not run the models again for this report.
What was measured
Stage R v8 is a Sarvam-M fine-tune in the local experiment record. The comparison uses the same suite manifest for the base and candidate in each task. Both record unsloth-4bit serving and think: false. The MMLU baseline’s serving and thinking fields were backfilled in its artifact; they are not contemporaneous metadata in that original file.
The frozen suite has six Indic languages—Hindi, Gujarati, Marathi, Bengali, Tamil and Telugu—in native and Romanized forms, plus English. Each direction has 100 GSM8K-Indic rows or 200 MMLU-Indic rows. “Romanized” means an Indic language represented in Latin script; it does not mean an English translation.
| Property | GSM8K-Indic | MMLU-Indic |
|---|---|---|
| Rows | 1,300 · 100 per direction | 2,600 · 200 per direction |
| Task | Generated numerical answer | Four-way multiple choice |
| Baseline | Sarvam-M, corrected KV-cache run | Sarvam-M, recorded base run |
| Candidate | Stage R v8 | Stage R v8 |
| Recorded mode | 4-bit · thinking disabled | 4-bit · thinking disabled |
The corrected GSM8K baseline matters. The comparison script selects the cached baseline file rather than the older base result that predates a serving fix. Mixing those files would attribute an inference-implementation change to training.
Two tasks, different gains
GSM8K-Indic and MMLU-Indic ask different questions. One requires generating a numerical result; the other compares candidate answer-letter likelihoods. Their accuracy values should not be averaged into a single “intelligence score.”
gsm8k-indic
mmlu-indic
Exact saved counts: GSM8K 963/1300 → 1092/1300; MMLU 1355/2600 → 1403/2600. Bars show observed accuracy, not confidence intervals.
View methods and limits
These bars use the checked-in Stage R v8 aggregate evidence, recounted from saved baseline and candidate rows. Accuracy is correct answers divided by evaluated rows. Both tasks use a 0–100% scale; they are not combined into a single score. This historical comparison does not measure NIF, retrieval or tool-use performance.
Scoring, setup and evidenceThe maths gain is substantial within this fixed evaluation. The knowledge-task gain is 1.85 percentage points. Neither measurement establishes factuality, conversational quality, safety, tool-use reliability or general transfer to an unfamiliar task.
The Romanized slices are not interchangeable
The aggregate hides different starting points. Romanized Marathi rose from 62% to 84%; Romanized Gujarati rose from 70% to 76%. Tamil improved from 56% to 66% and still had the lowest candidate accuracy among these six Romanized maths slices. A shared script label does not make the languages equally difficult for this model.
| Language | Base | Stage R v8 | Change |
|---|---|---|---|
| Hindi | 75% | 87% | +12 pp |
| Gujarati | 70% | 76% | +6 pp |
| Marathi | 62% | 84% | +22 pp |
| Bengali | 71% | 79% | +8 pp |
| Tamil | 56% | 66% | +10 pp |
| Telugu | 71% | 81% | +10 pp |
These are descriptive slice results. At 100 rows, one changed answer moves a slice by one percentage point. Item selection and relationships between translated questions matter before interpreting a small difference as a stable advantage.
The regression an average hides
On native-script Telugu MMLU-Indic, the base answered 104 of 200 rows correctly. Stage R v8 answered 95. Accuracy fell from 52.0% to 47.5%, a loss of 4.5 percentage points. The overall MMLU score still increased because other directions improved.
“The aggregate improved” is supported. “Every target language improved” is not.
A multilingual model should be inspected by language and script, not only by an average across directions. The regression deserves follow-up under controlled conditions; the present record does not establish its cause or prove that a later checkpoint fixes it.
How scoring works—and where it can mislead
The local runner scores MMLU-Indic by comparing the likelihoods of A, B, C and D at the answer position. That avoids treating an unparseable generated letter as a knowledge failure, but it is still a four-option task rather than an open-ended factuality evaluation.
GSM8K-Indic uses generation. The reference answer follows the dataset’s numerical answer marker; the model answer is extracted from its output. Formatting and truncation can affect the score. A passing numerical match is not an independent assessment of every intermediate reasoning step.
The comparison utility groups GSM8K answers by their original question when calculating paired statistics. It identified 840 original-question clusters among 1,300 rows here. Treating all translated or Romanized rows as independent observations would overstate the available evidence. We have not added invented error bars to the chart.
Compare the same items, not just two rounded scores
A percentage is a summary of item-level outcomes. To understand the difference between two checkpoints, match their answers on the same suite: which questions both answered correctly, which both missed, which the candidate gained and which it lost. Two models can have similar aggregate accuracy while disagreeing on many items. A gain is more informative when the comparison preserves those disagreements.
The saved GSM8K results contain 187 rows that the base missed and v8 answered correctly, alongside 58 rows that the base answered correctly and v8 missed. Their difference is 129 additional correct rows, matching the change from 963 to 1,092. For MMLU, the corresponding counts are 225 gains and 177 losses: a net increase of 48 correct rows, from 1,355 to 1,403. These are recorded paired outcomes, not a count of independent discoveries.
That last distinction matters. Native and Romanized versions can originate from the same underlying question. A model may gain both versions for a shared reason. The comparison utility groups related GSM8K rows by original-question identity when calculating paired statistics. We show the actual counts but do not turn every translated row into an independent sample or add confidence intervals without a fully specified protocol.
- ↓
Match the suite
Check item identity, sample counts and recorded configuration.
- ↓
Recount outcomes
Preserve gains, losses and language-level differences.
Limit the claim
Account for related questions and incomplete historical settings.
Method diagram. It describes the inspection of saved results, not a fresh model run.
Why publish a historical, deprecated checkpoint?
Stage R v8 is marked deprecated in the project’s model registry. It is not the current model recommendation, and this article is not an instruction to deploy it. We retain the report because the experiment has inspectable outcomes and illustrates a useful research lesson: a fine-tune can make a substantial gain on one task while producing a smaller gain elsewhere and a regression in one language direction.
A checkpoint’s status and its measured results answer different questions. The status helps an operator decide which version to investigate or use. The historical result records what a particular version did under particular conditions. Deprecating the version does not erase those observations, but publishing the observations must not make the version appear current.
Later development records exist. Their scores cannot simply be substituted into this chart: a new checkpoint may have a different training objective, and a valid comparison requires compatible evaluation settings and a review of the saved rows. We have not combined unrelated safety, serving or task-completion results into a single ranking. A model that improves numerical accuracy has not thereby established dependable refusal behavior, conversational helpfulness or coding-agent completion.
The purpose of this report is therefore explanatory and archival. Readers can inspect the arithmetic, see the language-level trade-offs and understand the limits of the method. A later release claim needs its own evidence bundle, rather than borrowing the strongest number from an older experiment.
What a follow-up experiment needs to establish
The native-script Telugu regression is a specific place to begin, not a diagnosis of the cause. A follow-up should first hold item identity and serving conditions constant, preserve the original baseline and candidate outputs, and inspect the questions on which the candidate lost a previously correct answer. The present counts cannot distinguish a training trade-off from sensitivity to answer positions or another evaluation condition.
For multiple-choice evaluation, answer-position changes offer one way to challenge a result. An experiment can rotate the options while preserving the question and correct answer, then report whether the model’s score depends on that ordering. That would be a new protocol with its own outputs and limitations. This article does not present an option-rotation result or infer one from the existing MMLU records.
For generated maths answers, retain the extraction rule and inspect formatting failures separately from a wrong numerical answer. Record output-token limits and truncation consistently. A new blind or held-out evaluation would also help separate development-suite progress from generalization. Its item selection and relationship to training material should be disclosed, not implied by the word “frozen.”
Before a broader release claim, evaluate the additional behaviors that matter for intended use. Helpful responses, harmful-request refusals, unnecessary refusals of safe requests and tool-driven task completion require different checks. They should have explicit denominators, failure examples and operating conditions. The two task-accuracy suites here cannot stand in for that work.
What this comparison cannot establish
A frozen manifest makes item identity checkable. It does not establish independence from all training data, a blind holdout, or immunity to checkpoint selection. The repository contains multiple development arms evaluated on these suites; repeated use can influence which run looks most attractive.
- Repeated evaluation across development arms can affect selection; this is not a new blind holdout.
- Translated and Romanized items are not all statistically independent.
- No confidence intervals or broad safety claim are inferred from these two accuracy suites.
- Historic output-token limits are not fully recorded in every result file; do not infer them from current runner defaults.
- MMLU baseline serving and thinking fields were backfilled in the stored artifact.
- Accuracy is task-specific; Telugu native-script MMLU-Indic regressed from 52.0% to 47.5%.
The historical generation limit is not fully recorded in every result. The current runner’s default must not be retroactively described as the exact setting used by every old run. The files are useful evidence with real limits, not a complete laboratory notebook.
A safety claim would require a separate evaluation, including harmful refusals and unnecessary refusals of safe requests. A capability-reuse claim would require a NIF experiment. Neither follows from the bars above.
Inspect the evidence
The downloadable summary contains aggregate scores, every saved language slice, source filenames and SHA-256 digests. It excludes prompts, raw model generations and local machine paths. The evidence-generation script verifies saved row counts and slice counts. Production builds validate the checked-in aggregate snapshot, so deployment does not require raw evaluation files.
Download the aggregate evidence JSON ↓aditya/eval/results/base_sarvam_m_cached_gsm8k-indic.jsonSHA-256 be9f3b87cfb491b23682beb4b8d1f71ea019184bd8ebea04c75a62e980920cab
aditya/eval/results/stage_r_v8_gsm8k-indic.jsonSHA-256 9953156a2891fcc5035c4364ef0d9573f2530de93ff1dda0255fb9012180023b
aditya/eval/results/base_sarvam_m_mmlu-indic.jsonSHA-256 00764c296af165dfd3e0f84d6a0407efbc6ba4a7f15bb5584278a25757cd437d
aditya/eval/results/stage_r_v8_mmlu-indic.jsonSHA-256 3cc5cef8c841fff3a2e64ea2c33a7ab688467c66ff225946f438d972e995d7c8
Inside the project checkout, recalculate the paired comparison without launching inference:
python aditya/eval/compare_arms.py stage_r_v8The methodology is in aditya/eval/run_frozen.py, the suite inventory in aditya/eval/frozen/MANIFEST.json, and the comparison logic in aditya/eval/compare_arms.py. This download contains the aggregate summary, not model weights or raw generations.