Models, RAG, MCP and NIF answer different questions
A practical comparison of generation, retrieval, tool connections and governed reuse—without turning unlike components into a leaderboard.
A responsibility comparison, not a competitive benchmark. Only the explicitly labelled saved model results are quantitative measurements.
Begin with the job, not the name
A model, a retrieval workflow and a tool protocol are often discussed as if they were competing ways to build the same thing. They occupy different parts of a system. A model proposes language and actions. Retrieval brings selected material into context. A protocol connects the host to operations. The host still has to manage task state, authority and evidence.
NIF adds a research question about governed capability reuse: when should a useful combination become something the system can discover and rely on later? That is not the same question as which model has the highest accuracy on a fixed task. It also does not replace the retrieval or connection layer.
This comparison names the responsibility, the relevant change boundary and the kind of evidence to inspect for each component. It is intended to help readers diagnose where a failure occurred and what a result actually measured. It does not rank NIF above the other pieces or claim that any single layer solves the whole task.
Model
Generate- Provides
- Candidate language and actions
- Inspect
- Task-specific model outcomes
RAG
Retrieve- Provides
- Selected external material
- Inspect
- Retrieval and grounded-answer checks
MCP
Connect- Provides
- Tool requests and observations
- Inspect
- Protocol, host policy and task behavior
NIF
Govern reuse- Provides
- Scoped capability lifecycle
- Inspect
- Reuse, abstention and withdrawal experiments
Conceptual comparison. These components can work together; the cards do not represent comparable performance scores.
How to read this figure
This figure organizes the responsibilities and evidence discussed in this section. It is a conceptual explanation, not a measured runtime trace, a benchmark or a claim that every mechanism has been validated end to end. Read the source inventory and research limits alongside it.
Source inventory and boundariesThe model generates; its quality is task-specific
A language model generates from its input and learned parameters under a chosen inference configuration. It can propose an explanation, a tool call or a patch. Those proposals can be useful, but a plausible response does not establish that an external action happened. The surrounding application needs an observation from the executed operation.
Model evaluation should name the task, checkpoint, dataset, scoring rule and operating conditions. Numerical-answer accuracy and multiple-choice accuracy are different measurements. A change in quantization, thinking mode or output limits can change the result without demonstrating the effect of a training recipe. Keep the comparison’s conditions visible before interpreting the score.
Changing a model or adapter can affect many behaviors at once. Inspect relevant regressions rather than presenting the strongest improved metric as a universal advantage. The historical Stage R v8 comparison later in this note is a concrete example: the aggregate maths gain and the native-script Telugu knowledge regression can both be true.
Retrieval makes material available at answer time
Retrieval-augmented generation combines generation with selected external material. It can help a system use information that is not available in its model parameters or is more current than the training record. The original RAG paper linked below studies a particular approach to combining parametric and retrieved information; it is not a result for Aditya’s implementation.
A retrieval workflow still needs separate checks. Did the search find relevant material? Was that material current and attributable? Did the generated answer accurately use it? A relevant document can be retrieved while the answer adds an unsupported statement. An answer can also look fluent while the retrieval missed the document that would have contradicted it.
Updating documents or an index is different from updating model weights. The resulting system may answer some questions differently without a new model checkpoint. When reporting a comparison, disclose the corpus and retrieval configuration so that improved information access is not silently described as improved parametric knowledge.
MCP standardizes the connection to tools
The Model Context Protocol provides a way for an application to discover tools and exchange requests and results with connected servers. Its tools specification describes the interface and recommends visible activity and human control over invocations. The host application remains responsible for deciding which connections to expose and how to present their authority boundaries.
A schema-valid call establishes something about the shape of a request, not the appropriateness of the action. A server response establishes an observation, not completion of the whole task. A connected write-capable tool can change external state, so its availability must not be treated as general permission to use it.
Inspect the host’s configured servers, permissions, approval behavior and handling of returned material. A standard protocol can improve interoperability while leaving application-level correctness and security questions open. The protocol should be evaluated at its own level, with task outcomes checked separately.
The host joins the pieces without inheriting their guarantees
The host tracks the task and coordinates model proposals, retrieved material and tool observations. It supplies the policy and completion logic that a connection protocol alone does not define. In Aditya, these responsibilities are represented in the runtime’s state, capability metadata, approval and recorded-verification mechanisms.
Integration introduces its own failure cases. A correct tool result can be attached to the wrong task. A past failed check can disappear from the summary. An approval can remain pending while the final answer claims the operation completed. A retrieved instruction can be mistaken for an authoritative request. Testing each component separately does not rule out these interactions.
A useful end-to-end evaluation follows the actual run. Preserve what the person requested, which operations were proposed and allowed, what executed and which check supports the final claim. The unit of inspection is the task outcome and its record, not simply the number of connected components.
- 01
Request
Define the desired behavior and scope.
- 02
Proposal & policy
Choose an operation and decide whether it may run.
- 03
Observation
Keep the result, including a failed or blocked operation.
- 04
Verification
Check the outcome before accepting completion.
A failed check can return the task to investigation. This is a method diagram, not a measured live trace.
How to read this figure
This figure organizes the responsibilities and evidence discussed in this section. It is a conceptual explanation, not a measured runtime trace, a benchmark or a claim that every mechanism has been validated end to end. Read the source inventory and research limits alongside it.
Source inventory and boundariesNIF investigates what should be retained for later work
NIF’s intended lifecycle addresses reusable capability after an attempt has been inspected. It asks how a procedure, memory or compiled adapter can remain discoverable, scoped, verifiable and withdrawable. That lifecycle can use models, retrieval and tools; it is not a substitute implementation of those components.
The current registry performs symbolic metadata discovery. The adapter bank records versions and activation state. Those are concrete mechanisms to inspect. A learned router, dependable transfer to unfamiliar tasks and successful automatic consolidation require different experiments. A module name or architecture diagram cannot supply the missing result.
Compare NIF against a baseline relevant to reuse, not against MCP as if a protocol were an agent. Hold task conditions constant, inspect first use and reuse, include cases where activation should be declined, and test withdrawal. The evidence should show whether retention changes useful outcomes rather than merely reducing the effort of finding a candidate.
Follow one illustrative investigation across the boundaries
Suppose an operator asks why a project’s session recovery test is failing. The model proposes inspecting relevant code. Retrieval or search helps locate material, and a file tool returns an observation. The host records that observation and applies policy before a proposed write. A test later checks whether the changed behavior satisfies the acceptance criterion.
Each component can work at its own level while the task still fails. Search may find the right file but the model may misread it. The write may execute successfully but alter the wrong condition. The test may pass without reaching the reported failure. The host’s final answer needs to identify the evidence actually collected, not combine these partial successes into an unsupported completion statement.
This is an illustrative walkthrough, not a reported successful repair. A later NIF retention decision would require another inspection: what can be reused, under which conditions, and what check supports that scope? Solving the present task and justifying future reuse are different decisions.
Inspect a local file
Check relevance and workspace scope before treating the result as task evidence.
Change project state
Confirm the requested target and apply the operation’s current approval policy.
Publish externally
Separate publication authority from a locally passing build or test.
Examples of distinct decisions, not an exhaustive policy matrix or security certification.
How to read this figure
This figure organizes the responsibilities and evidence discussed in this section. It is a conceptual explanation, not a measured runtime trace, a benchmark or a claim that every mechanism has been validated end to end. Read the source inventory and research limits alongside it.
Source inventory and boundariesName what changed before naming what improved
A system can change because of a model checkpoint, an adapter, a document corpus, a retrieval index, a tool server or host policy. Those changes have different effects and different reproducibility requirements. If several change together, the outcome may be useful without identifying which one caused the improvement.
Keep a comparison’s versions and configuration explicit. If new documents supply the answer, describe the contribution as information access rather than silently crediting a model update. If a tool server becomes available, distinguish environment availability from model competence. If a host gate rejects more tasks, inspect whether it rejects incorrect completions or merely blocks useful work.
The evidence should make it possible to reproduce the changed boundary. Source digests identify files; they do not explain the experiment on their own. Include the relevant setup and exclusions so that an exact identifier does not hide an incomplete methodology.
- Model behavior
- Checkpoint, adapter and inference configuration
- Information access
- Documents, index and retrieval configuration
- Available operations
- Connected servers, tool versions and permissions
- Task decisions
- Host policy, acceptance criteria and completion rules
A version checklist, not a causal attribution of an observed improvement.
How to read this figure
This figure organizes the responsibilities and evidence discussed in this section. It is a conceptual explanation, not a measured runtime trace, a benchmark or a claim that every mechanism has been validated end to end. Read the source inventory and research limits alongside it.
Source inventory and boundariesChoose a check that can reject the specific claim
For a model claim, inspect task-level performance under compatible conditions. For retrieval, inspect whether relevant evidence was found and used accurately. For a connection layer, inspect discovery, request and result behavior. For an agent task, inspect the executed work and acceptance criterion. For governed reuse, inspect activation, abstention, transfer and withdrawal.
Do not treat passing one of these checks as passing all of them. A connector can return a valid result while the model draws a wrong conclusion. A model can score well on a fixed numerical suite while a host fails to preserve approval state. A reuse registry can find a candidate that should not be activated.
Report failures and inadmissible cases explicitly. An environment preventing execution is not the same observation as an executed check rejecting the attempted solution. The denominator should reflect the protocol, not whichever subset makes the system look strongest. This is how a comparison remains useful to a reader trying to understand where the evidence ends.
A real result still belongs to its measured task
The saved Stage R v8 report contains a model comparison, not a comparison among Model, RAG, MCP and NIF. On 1,300 GSM8K-Indic rows, the corrected Sarvam-M base answered 963 correctly and v8 answered 1,092. On 2,600 MMLU-Indic rows, the counts were 1,355 and 1,403. The chart below uses the same inspected aggregate snapshot as the full report.
The maths improvement is much larger than the knowledge-task change. Native-script Telugu MMLU accuracy fell from 52.0% to 47.5% even though the overall MMLU score increased. That is why a task and language breakdown is more informative than a single system-wide score.
Stage R v8 is deprecated in the project registry, and these are historical saved evaluations rather than new inference. They do not measure retrieval quality, MCP interoperability, coding-agent completion or NIF reuse. The full report explains the corrected baseline, related question clusters and incomplete historical settings before drawing a bounded conclusion.
gsm8k-indic
963/1300 → 1092/1300 correct
mmlu-indic
1355/2600 → 1403/2600 correct
Observed saved accuracy, without invented confidence intervals. Read the setup and limitations →
View methods and limits
These bars use the checked-in Stage R v8 aggregate evidence, recounted from saved baseline and candidate rows. Accuracy is correct answers divided by evaluated rows. Both tasks use a 0–100% scale; they are not combined into a single score. This historical comparison does not measure NIF, retrieval or tool-use performance.
Scoring, setup and evidenceWhat this comparison does—and does not—conclude
These components are often complementary. Choose them by the responsibility the application needs, then evaluate the resulting behavior at the right level. None supplies a universal guarantee of correctness, safe action or useful retention. In particular, this note does not establish a performance advantage for NIF over the other components.
For background, read the original RAG paper and the official MCP tools specification. For Aditya’s implementation boundary, inspect the capability registry and completion code. For a measured model result, read the full saved-evaluation report rather than inferring an experiment from the conceptual diagrams. Those sources answer different questions, and this comparison keeps that difference visible.
Sources and related evidence
Primary research paper, 2020; background, not an Aditya result
Official connection-layer specification
agent_system/kernel/registry.py; completion.py
Inspected aggregate counts and source digests