From a model proposal to a checked operation
How interfaces, run state, context, capability discovery, policy and verification share responsibility in Aditya’s agent runtime.
A source-grounded architecture note. Diagrams explain responsibilities, not a literal deployment topology or a reliability benchmark.
A model response is not the whole system
A model can suggest a useful action without knowing whether that action was executed. It can also produce a convincing completion statement while a test is still failing. An agent runtime needs sources of truth outside that statement: the requested task, the policy decision, the executed operation and the recorded observations.
Aditya separates those responsibilities so that an operator can inspect the run rather than infer it from the conversation. The model proposes language and actions. The runtime tracks state, makes operations available under policy, preserves results and evaluates whether a completion claim has support. This separation creates opportunities to challenge an incorrect proposal; it does not make the system infallible.
The map below describes these responsibilities rather than a strict stack of isolated services. Policy and verification apply across a run. Context depends on earlier observations. Interfaces expose selected parts of the record. Treat the diagram as a reading aid for the implementation, not evidence of a particular hosting topology or performance result.
- 01
Interface
Expose task activity, approvals and results.
- 02
Run state & context
Preserve commitments, observations and relevant history.
- 03
Capability discovery
Find candidate tools, skills and registered units.
- 04
Policy & verification
Decide what may run and what the evidence supports.
- 05
Foundation model
Propose language and actions under the configured model.
Responsibilities overlap: policy applies to operations and verification reads the record. This is not a literal service stack.
How to read this figure
This figure organizes the responsibilities and evidence discussed in this section. It is a conceptual explanation, not a measured runtime trace, a benchmark or a claim that every mechanism has been validated end to end. Read the source inventory and research limits alongside it.
Source inventory and boundariesDifferent interfaces expose different entry paths
The web and desktop workspaces expose server-backed task activity. The Python CLI can invoke the agent loop directly from a chosen workspace. These are different entry paths into project work, not interchangeable claims about where every component runs. The configured model endpoint, connected tools and enabled policy determine what an actual deployment can reach.
An interface should make useful distinctions visible: a proposed action versus an executed tool call, a pending approval versus an allowed operation, and an attempted fix versus a checked outcome. A compact terminal can present those distinctions differently from a desktop inspector, but neither should convert a missing result into a success indication.
The public model page contains interface illustrations. Their project names and activity are examples, not live backend sessions. For the current setup and entry points, use the documentation. For model-quality measurements, use a report that names the checkpoint, suite and evaluation conditions; a polished interface is not such a measurement.
Run state keeps commitments outside the narration
A task record needs to preserve what the operator requested and which conditions count as acceptance. It also needs to distinguish planned work, running operations, blocked steps and completed work. If those commitments exist only in the model’s latest paragraph, a fluent response can accidentally replace an unfinished plan.
The current kernel represents run state and recorded verification independently of the final answer. Completion logic can inspect open criteria, unfinished plan steps and unresolved approvals. It also checks for structurally invalid plans, such as cycles or dangling dependencies, so that ‘nothing runnable’ is not automatically treated as ‘everything finished.’
The mechanism still depends on what was recorded. A missing or poorly chosen acceptance criterion can weaken the task definition. A passing check that does not reach the changed behavior can weaken the evidence. The architecture supplies a place to preserve and inspect those decisions; it cannot make an inadequate criterion meaningful by storing it.
- 01
Request
Define the desired behavior and scope.
- 02
Proposal & policy
Choose an operation and decide whether it may run.
- 03
Observation
Keep the result, including a failed or blocked operation.
- 04
Verification
Check the outcome before accepting completion.
A failed check can return the task to investigation. This is a method diagram, not a measured live trace.
How to read this figure
This figure organizes the responsibilities and evidence discussed in this section. It is a conceptual explanation, not a measured runtime trace, a benchmark or a claim that every mechanism has been validated end to end. Read the source inventory and research limits alongside it.
Source inventory and boundariesContext should preserve contradictions as well as clues
The model’s next decision depends on the material available in context. Useful context includes the task, relevant project information and earlier observations. It should also preserve evidence that contradicts the current explanation. Reading a file is not progress if the run continues to ignore the behavior that file actually implements.
Long investigations create a practical tension: enough history must remain to support the decision, but not every unrelated event belongs in every prompt. Summaries need to preserve failed checks, changed assumptions and unresolved conditions. Dropping those details can make a repeated failed approach look like a fresh plausible idea.
Memory and retrieved material add another boundary. A past observation may be useful without being current, and a retrieved document may contain instructions without becoming authoritative. Compare its source and assumptions with the present task. The current architecture does not justify a blanket claim that every context assembly is complete, unbiased or resistant to malicious input.
Capability discovery is a metadata operation
The registry provides a shared description of available native tools, configured MCP tools, skills and registered capability units. Discovery ranks candidates using symbolic metadata matching. It does not train a router, execute the returned operation or establish that a candidate is appropriate for the user’s request.
Inspecting a candidate should reveal its provider, declared permissions, risk metadata, version and availability. The actual arguments still matter. A shell tool can be used for a harmless read or a consequential write; its static label cannot settle the risk of every command. Apply the operation policy to the specific proposal rather than relying only on the tool’s name.
This separation keeps connection, selection and execution from collapsing into one step. A tool can be discoverable but unavailable at execution time. A relevant capability can still require approval. A healthy connector can return material unrelated to the question. Each observation should remain visible so that the next decision can respond to it.
Permission is checked before execution, not after success
Correctness and authorization are different questions. A model might propose a technically effective change outside the requested files, or a well-tested deployment that the operator has not approved. Successful verification does not retroactively authorize those actions. The policy boundary needs to apply before the consequential operation.
For a project task, inspect the exact target and scope. Reading a local file differs from sending it to a remote service. Running a test differs from publishing a release. A directory chosen as a workspace is useful context, but it should not be advertised as complete host isolation. Optional tools and external connections can introduce their own data and permission paths.
The runtime has approval and policy mechanisms to represent these decisions. Operators still need to understand their configuration. A confirmation prompt is an opportunity to decide, not a certificate that the approved action is harmless. The conceptual cases below show why a tool’s availability and the task’s authorization must remain separate.
Inspect a local file
Check relevance and workspace scope before treating the result as task evidence.
Change project state
Confirm the requested target and apply the operation’s current approval policy.
Publish externally
Separate publication authority from a locally passing build or test.
Examples of distinct decisions, not an exhaustive policy matrix or security certification.
How to read this figure
This figure organizes the responsibilities and evidence discussed in this section. It is a conceptual explanation, not a measured runtime trace, a benchmark or a claim that every mechanism has been validated end to end. Read the source inventory and research limits alongside it.
Source inventory and boundariesExecution returns an observation, not a verdict
A tool result can establish that an operation returned certain output. What that output means for the task is a further question. A search result may narrow an investigation. A compiler error may reject the proposed change. A command that exits successfully may still test an unrelated behavior. Preserve the output before drawing the conclusion.
The record should distinguish the proposed call from the operation that actually ran, including errors, unavailable dependencies and denied actions. If a tool never executed, its expected output cannot appear as observed evidence. If a service returned an ambiguous result, do not force it into a success or failure category without inspecting what is known.
External writes deserve particular care. A timeout can occur after a service accepted an operation, so an immediate retry may duplicate the effect. Use a read-only status check or operation identifier when available. If state cannot be determined, preserve the uncertainty and stop at the appropriate authority boundary.
Recovery should revise the strategy
Failure can be informative when it changes the explanation. An assertion rejecting a patch is different from a test runner failing during collection. A connector unavailable in the environment is different from a tool returning an empty result. Those conditions require different next steps, and they should not be pooled into a generic failure narrative.
A useful recovery loop inspects the observation, narrows the hypothesis and selects another permitted check. Repeating an unchanged command can test whether a fault was transient, but persistent identical failure is not progress by itself. Nor does switching to a more powerful tool resolve missing authorization.
Blocked and unverified states are legitimate results. If a necessary dependency, service or approval is unavailable, identify what remains and why. A run should not broaden its task or narrate success simply to avoid admitting that the required check could not be completed. This is a design requirement, not a claim that every current model run always follows it.
The completion gate challenges the final sentence
The model can propose that a task is complete. The completion gate inspects recorded conditions instead of accepting that proposal solely because it is fluent. Open acceptance criteria, unresolved approvals, unfinished work and the most recent failed verification can supply reasons not to finish.
The code also contains checks against vacuous completion: changes with no verification, or change tasks with no meaningful investigation. Those mechanisms address specific ways a run can look finished without supporting the requested outcome. Their existence is evidence of implementation, not a measured guarantee that no false completion is possible.
Verification coverage remains important. A test can pass on the original broken behavior, and a general suite can miss the particular regression being repaired. The handoff should identify which check exercised the acceptance criterion and what remains outside it. Passing a runtime gate does not justify a broader statement such as ‘production-ready’ without additional release evidence.
Evaluate mechanisms and model behavior at different levels
Unit tests can exercise a policy rule, an event record or a completion decision using fixtures. End-to-end tasks can inspect whether a model-driven run uses those mechanisms effectively. A live serving check can inspect another boundary again. These layers are complementary; a passing fixture-based test is not a task-success percentage.
For an agent experiment, report how many tasks were attempted, how many were admissible for grading and which were excluded by infrastructure or verifier faults. Preserve the distinction between an assertion rejecting the model’s work and an environment preventing a valid test. Keep denominators explicit so that a clean subset does not quietly stand in for the full workload.
Model language evaluations also answer a different question. The Stage R v8 report measures saved numerical and multiple-choice answers. It does not measure whether this runtime completed engineering tasks safely or whether the desktop interface was usable. Read each result at the level where its protocol collected evidence.
Inspect the source and the deployment boundary
Start with run state, capability discovery, approval and completion. Then follow the CLI or server path appropriate to the interface you are using. Configuration matters: model endpoints, enabled tools, workspace access and external services can change the behavior and data paths of a deployment without changing the conceptual diagram.
This note deliberately does not publish secrets, machine paths or pretend that a local source inspection certifies every deployment. The documentation describes the current setup boundaries. The journal explains verification in practical terms. NIF extends the research question toward governed reuse, but the runtime map alone does not demonstrate that broader thesis.
Sources and related evidence
agent_system/kernel/run_state.py; events.py
agent_system/kernel/approval.py
agent_system/kernel/completion.py; plan_dag.py
agent_system/cli.py; scripts/agent-server.sh