NIF: when should useful work become reusable capability?
The proposed capability lifecycle, its current runtime foundations, and the experiments needed to establish dependable reuse.
Architecture under investigation. Implementation mechanisms are described separately from unproved reuse and transfer claims.
The question starts after a successful attempt
A system repairs a configuration and the relevant check passes. That is useful work. The harder question comes next: should the system keep a procedure from that repair and rely on it in another project? The original result may depend on a particular dependency version, environment or permission. Reusing it without those conditions can preserve a mistake as easily as a solution.
Neural Intelligence Fabric is Aditya’s research architecture for that decision. Its intended lifecycle is to discover, compose, execute, verify and then retain or reject scoped capability. It sits around a foundation model rather than replacing it. The goal is not to remember every output or let every conversation alter the model. It is to keep an inspectable connection between a reusable candidate and the evidence that supports it.
This note explains the responsibilities in that lifecycle, identifies mechanisms in the current source, and sets out what a convincing evaluation would need. It does not report a measured NIF transfer advantage. The distinction matters: implemented bookkeeping makes an experiment possible, but the experiment still has to show whether reuse helps.
A procedure, a memory and an adapter preserve different things
A written skill preserves a method. A memory entry preserves information from earlier work. A compiled adapter preserves a versioned change in model behavior. All three can contribute to a task, but they fail in different ways. A procedure can have stale prerequisites; a memory can contain a superseded conclusion; an adapter can improve one task while reducing accuracy on another.
Calling these objects capabilities is useful only if their differences remain visible. The relevant question is what each object does, what conditions it requires and what evidence justified its use. A previous successful command is not automatic permission to execute it again. A stored adapter version is not proof that activating it will improve the present task.
The intended record should therefore distinguish the reusable object from the observation that led to it. ‘This test passed on this revision’ is narrower than ‘this procedure works across repositories.’ Keeping those statements separate gives a reviewer a way to accept the observation without accepting an unsupported general rule.
- Origin
- The task, source and version that produced the candidate.
- Preconditions
- The environment and dependencies its earlier check covered.
- Authority
- The current request and permission for the specific operation.
- Observation
- What actually ran and what the operation returned.
- Acceptance
- The check that can support or reject the outcome.
- Withdrawal
- The conditions that should narrow or invalidate reuse.
Record requirements, not a claim that every current object contains complete evidence.
How to read this figure
This figure organizes the responsibilities and evidence discussed in this section. It is a conceptual explanation, not a measured runtime trace, a benchmark or a claim that every mechanism has been validated end to end. Read the source inventory and research limits alongside it.
Source inventory and boundariesDiscovery finds candidates; it does not select a winner
The current capability registry unifies metadata from native tools, configured MCP tools, skills and registered units. It ranks candidates using keyword and metadata overlap. This is symbolic discovery, not a learned router over neural capability embeddings. A relevant-looking result is a candidate for inspection, not an established best action.
That boundary is useful in practice. Discovery can narrow the material presented to the model while leaving suitability and authorization open. Inspect the candidate’s provider, version, health, trust, permissions and schema. Then compare those fields with the actual operation being considered. Static tool metadata cannot fully describe the risk of a shell command whose arguments have not yet been chosen.
Registration and retrieval do not themselves execute the capability. Execution needs its own policy decision and result record. A future learned router would require its own evaluation of selection quality, unsuitable activation and behavior on unfamiliar tasks. It cannot be inferred from the existence of today’s metadata index.
Compose around the task’s present conditions
A saved method may be appropriate for one repository and wrong for a superficially similar one. Before composing a run, compare the present task with the candidate’s prerequisites. Which files and tools are available? Which dependency versions matter? What did the operator ask to change, and what did they ask only to inspect? A description that matches the vocabulary of the task can still miss its cause.
Relevant context should carry both supporting and contradicting evidence. If a previous procedure assumes a local database but the current project uses a remote service, that mismatch belongs in the decision. If the current request excludes writes, a stored repair procedure cannot widen that authority. Composition is not merely choosing pieces that can technically connect; it also means checking whether they belong in this run.
The full research question is whether a system can make these choices dependably. The current mechanisms allow scope and observations to be represented, but this note does not claim automatic suitability assessment is solved. Changed environments and narrowed permissions are important negative cases for a later evaluation.
- 01
Discover
Find an existing candidate.
- 02
Inspect & compose
Match scope to the current task.
- 03
Execute
Apply policy and preserve observations.
- 04
Verify
Exercise the acceptance criterion.
- 05
Retain or reject
Keep only a justified, bounded record.
A failed check can return the task to investigation. This is a method diagram, not a measured live trace.
How to read this figure
This figure organizes the responsibilities and evidence discussed in this section. It is a conceptual explanation, not a measured runtime trace, a benchmark or a claim that every mechanism has been validated end to end. Read the source inventory and research limits alongside it.
Source inventory and boundariesPreserve what happened, including the failed attempt
A proposed operation, an executed operation and a successful outcome are different events. The model may propose reading a file that does not exist. A command may finish without testing the behavior the operator cares about. A service may time out after accepting a write. The record needs to preserve the actual observation rather than replacing it with the expected result.
Failures are useful when they can change the next decision. An unavailable tool may require another source or a clear blocker. A failed assertion may challenge the proposed repair. A timeout on a consequential action should prompt inspection of state before retrying. Repeating the same operation without using its result does not turn a long trace into progress.
The example of a configuration repair in this note is illustrative, not an unpublished successful NIF run. In a real experiment, preserve the input conditions, allowed operation, observations and verifier outcome. Without that record, a later retention decision rests on a story about the attempt rather than evidence from it.
Verification must be able to reject a convincing result
A useful verifier checks the requested outcome, not the fluency of the explanation. For a repair, a regression test should reach the broken path and be capable of failing when that path is still wrong. A process exit code can support a narrower observation—such as a command completing—but it does not automatically establish that the task was satisfied.
The current completion logic inspects recorded verification, open acceptance criteria, unfinished work and unresolved approvals. It can challenge a completion proposal when the most recent check failed or a change lacks verification. These are runtime mechanisms. Their effectiveness still depends on the chosen criterion, the quality of the check and the evidence stored in the run.
A passing local check can justify a bounded conclusion. It cannot establish that the procedure transfers to every environment, or that retaining it is harmless. Verification of the first attempt and evaluation of later reuse are separate steps; the latter must include cases that the first test never exercised.
Retention should add a boundary, not a permanent privilege
Once a run is checked, the next decision is what to retain. Sometimes the useful result is an observation, not an executable procedure. Sometimes a procedure should remain a draft because its prerequisites cannot be established. Rejection is a meaningful outcome: a system that keeps every candidate has accumulated material, not demonstrated governed reuse.
A proposed retention record should include origin, supported scope, relevant versions, the acceptance check and reasons to withdraw it. These are inspection requirements, not a claim that every current object contains a complete laboratory record. New evidence should be able to narrow the scope rather than silently inherit the broadest previous interpretation.
Approval for one operation also must not become standing approval for future operations. A later task may involve a different project, different people or an external service. Apply the present permission policy again. Reuse may save investigation effort; it should not remove the decisions needed to authorize the action.
The adapter ledger records versions, not task superiority
The reversible adapter bank tracks compiled versions, base-model compatibility, declared rank, measured norm metadata, activation and rollback history. It is a ledger, not a model loader. Tensor compilation and measurement occur in separate code. Recording a norm policy decision is different from proving that the resulting behavior is useful on the task.
An adapter should be evaluated in isolation and, where composition is intended, in the combinations that will be used. A measured improvement for one unit does not establish that combining it with another preserves that improvement. Compatibility metadata helps identify conditions to inspect; it is not a substitute for a composed-behavior evaluation.
The present source supports concrete questions about version identity and activation state. Broader questions about transfer, interference and consolidation remain experimental. We keep those statuses visible so that a diagram of the architecture cannot be mistaken for a reported capability result.
Mechanisms in the source
- Symbolic capability discovery
- Recorded state, policy and completion checks
- Adapter version and activation ledger
Questions requiring experiments
- Reliable selection on unfamiliar tasks
- Useful transfer and composed behavior
- Dependable withdrawal and consolidation
A qualitative status summary from inspected code. There is no readiness percentage or measured transfer score implied.
How to read this figure
This figure organizes the responsibilities and evidence discussed in this section. It is a conceptual explanation, not a measured runtime trace, a benchmark or a claim that every mechanism has been validated end to end. Read the source inventory and research limits alongside it.
Source inventory and boundariesWithdrawal needs a test of its own
Disabling a procedure, withdrawing a memory entry and deactivating an adapter are different operations. Each can affect dependent records or active runs. A rollback ledger can identify the intended previous version without proving that every surrounding system has returned to its previous behavior. The test must exercise the relevant state, not only confirm that a history entry was appended.
For a reuse experiment, include a capability later found to be unsuitable. Check whether new tasks stop selecting it, whether affected records remain inspectable, and what happens to a run already using it. Also inspect unrelated capability: withdrawing one candidate should not be described as clean merely because its own flag changed.
These are proposed evaluation cases, not reported successful rollback trials. They make reversibility concrete. The goal is an operator who can see what was withdrawn, why, and which effects remain uncertain. A silent switch to another candidate would hide the evidence that should have changed the system’s trust.
Test reuse, transfer and abstention separately
Begin with a controlled first-use versus reuse comparison on the same task conditions. Record correctness, inspection cost and recovery from failure separately. Faster discovery is useful, but it is not the same result as better task completion. A run that avoids one search while producing a wrong repair has not established a reuse advantage.
Then test changed conditions and unfamiliar tasks. Include cases where the correct decision is not to activate the saved candidate: a missing prerequisite, a different cause, a changed dependency or a narrower request. Record unsuitable activation and correct abstention alongside successful reuse. Otherwise the evaluation rewards selection without examining whether the selected capability belonged in the task.
Finally, challenge withdrawal and composition. Use explicit baselines and preserve unsuccessful attempts. The language-accuracy results elsewhere on this site do not answer these questions merely because they belong to Aditya. A NIF claim needs a NIF experiment with its own task definition, denominator and evidence.
Current limits and inspectable sources
NIF is a research direction with real implementation foundations. Symbolic discovery, run state, policy, completion checks and adapter bookkeeping can be inspected in the source. A learned router, dependable general transfer and automatic consolidation are not established by those modules. This note supplies a map of the work and the experiments still required, not a claim that the thesis has been proved.
Read the registry and adapter-bank module descriptions alongside the completion code. They state important scope boundaries, including what metadata retrieval does and what the adapter ledger does not load or measure. The accompanying journal note explains the same concept in a shorter educational format. The separate Stage R v8 report contains measured language-task results, with its own limitations.
Sources and related evidence
agent_system/kernel/registry.py — symbolic metadata retrieval
agent_system/kernel/adapter_bank.py — version and activation ledger
agent_system/kernel/completion.py — evidence-based completion checks