ARCHITECTURE
Deterministic Agent Architecture: GraphRAG Memory, Human-in-the-Loop and Isolated Tool Execution
Most enterprise agent projects fail for the same three reasons: the model cannot reach the systems that hold the answer, nobody can predict or constrain what it will do, and nothing it did can be reconstructed afterwards. Bi-Mind is built around those three constraints rather than around the model.
Deterministic orchestration instead of an open-ended loop
A single general-purpose agent with a large tool belt is non-deterministic by construction: the same request can take a different path on every run, and the blast radius of a wrong path is the union of every tool it holds. Bi-Mind splits that into an orchestrator and a set of specialist agents. The orchestrator's only job is to decide which specialist owns the request; the specialist holds a small, reviewed tool set scoped to its domain.
The consequence is operational rather than academic. A specialist can be reasoned about, tested and signed off by the team that owns that domain. Adding a capability means adding or changing one specialist, not re-validating a monolithic prompt. And when something goes wrong, the question "which agent did this and what was it allowed to do" has one answer.
- Bounded tool surface: Each specialist is given only the tools its domain requires, so the set of possible actions is enumerable and reviewable.
- Routing is a decision, not a side effect: The orchestrator chooses a specialist and records that choice; it does not itself touch enterprise systems.
- Workers for long work: Work that outlives a chat turn is handed to a background worker that streams progress and completion back into the same conversation.
GraphRAG hybrid memory
Plain vector retrieval answers "what text looks like this question". Enterprise questions are usually relational: which contract covers this site, which change preceded this incident, which part belongs to which assembly. Flat similarity search returns plausible passages and loses the structure that made the answer correct.
Bi-Mind combines a graph layer, which holds entities and the relationships between them, with vector retrieval over document content and a key-value working memory for the live conversation. A query can walk the graph to find the relevant entities, then retrieve the passages attached to them, instead of hoping that one embedding happens to capture both.
The practical effect is fewer confident-but-wrong answers on exactly the questions enterprises care about — the ones whose answer depends on a relationship rather than on a phrase.
- Graph layer: Entities and their relationships, so a question can be resolved by traversal rather than by similarity alone.
- Vector retrieval: Content-level search over indexed documents, scoped to what the asking user is allowed to see.
- Working memory: Conversation state that survives across turns and, where configured, across sessions, so follow-up questions do not restart from zero.
Isolated tool execution and MCP
Tool calling is where an agent stops being a text generator and starts being an actor, so it is the layer that decides whether the platform is safe to deploy. In Bi-Mind the model never executes anything itself. It emits an intent; a separate execution layer validates that intent against the caller's permissions and the tool's declared contract, runs it, and returns a structured result.
Connecting enterprise capability happens through the Model Context Protocol (MCP) and equivalent declared interfaces, which means a tool is a contract with a schema rather than an ad-hoc prompt convention. A tool the caller's role does not grant is not merely discouraged — it is not reachable.
Separating intent from execution is what makes the rest tractable: permissions are enforced in one place, every call has a recorded shape, and a misbehaving model produces a rejected call rather than an unexpected write.
- Intent and execution are separate: The model proposes; a separate layer authorizes and runs. A model cannot widen its own reach.
- Permissions follow the person: A tool call runs with the signed-in user's permissions under the active tenant, never with an ambient service identity.
- Declared contracts: Tools are described by schema, so arguments are validated before anything executes.
The human approval gate
Read operations and side-effect operations are not the same risk and are not treated the same way. Any action that changes a system of record, moves money, writes a configuration or sends something outward can be declared as requiring approval. The agent then does the analysis, states exactly what it is about to do, and stops.
An authorized person approves or rejects in the same conversation, and only then does execution proceed. This is the mechanism that lets an organization put an agent in front of a production system at all: the autonomy is in the reasoning, the authority stays with a person.
The gate is channel-independent. A request raised from the web workspace, from an assistant embedded in another application, or from a chat on a phone goes through the same approval step.
Tenant isolation and zero data leakage
Every request carries an active tenant, and data access is scoped to it end to end — retrieval, tool execution and conversation history alike. Role-based access control sits on top: roles decide which agents a person may use, which they may manage, and which tools are reachable at all.
For organizations that cannot let data leave the perimeter, the platform runs on-premise, including fully air-gapped installations with locally hosted models. The deployment shape changes; the permission model, the approval gate and the audit trail do not.
End-to-end audit trail
Every run is recorded as a sequence: the request, the routing decision, each tool call with its arguments and result, each approval with the person who gave it, and the final answer. The point is not compliance theatre — it is that an agent whose actions cannot be reconstructed cannot be debugged, improved or defended after an incident.
This is also what makes a deterministic architecture worth having. Because the path is bounded, the recorded trace is readable: a reviewer can follow what happened without reverse-engineering an open-ended loop.
FAQ
What does "deterministic" mean for an AI agent?
It does not mean the language model is deterministic. It means the set of actions available at each step is bounded and known: an orchestrator routes to a specialist, the specialist holds a small reviewed tool set, every tool call is authorized against the caller's permissions, and side-effect actions require explicit human approval. The reasoning varies; the action space does not.
Why GraphRAG instead of plain vector RAG?
Because enterprise answers usually depend on relationships — which contract covers this site, which change preceded this incident — and flat similarity search loses that structure. A graph layer lets a query resolve the right entities by traversal and then retrieve the passages attached to them, which removes a large class of confident-but-wrong answers.
How is data kept from leaking between tenants or out of the company?
Every request carries an active tenant and all retrieval, tool execution and history are scoped to it; tool calls run with the signed-in user's permissions rather than a shared service identity. For stricter requirements the platform runs on-premise or fully air-gapped with locally hosted models, so no content has to leave the perimeter.
Can an agent act without a human?
Read and analysis work, yes. Actions that change a system of record, move money, write configuration or send something outward can be declared as requiring approval; the agent prepares the action, states its effect and waits for an authorized person to approve it in the conversation.
Related
- Energy & SCADA — Agentic AI for Energy and SCADA Operations
- ERP & Finance — Agentic Reconciliation and Operations for ERP
- Database & IT Ops — Agentic AI for Database and IT Operations
Bi-Mind · Privacy Policy · Terms of Service · info@bi-mind.com