Where No Agent Has Gone Before
For thirty years, computing has had one server side and one client side. The server had a fixed address, a published interface, trained operators, monitoring and a billing record. The client did none of that work: it opened a connection, sent a request and consumed the response. The server lived in the data center; the clients were your phone, your car, the camera above the loading dock, and the whole discipline of running a dependable system lived on one side of the connection.
Now the topology changes. Every device can serve its neighbors: one runs a model, one holds a cache, one relays to the network. The mesh is far more capable than the client-server arrangement was, but every device has quietly become a service that other components depend on. Who authenticates the caller? Who meters the work? What happens when the request fails halfway through?
That change is happening to devices right now, and it is the story of this post. AI agents are about to run on the phone, in the vehicle, behind the camera and on the factory PC, not only in distant data centers. The moment a device starts doing work for someone else, it inherits every obligation the server carried. mimOE is how mimik puts the agent on the device and brings those obligations along with it, in one package. Here is why that matters, who it matters to, and what it looks like when it works.
The imbalance nobody designed
The client-server arrangement was never a decision. It was a consequence of topology. The server had a stable address, a power supply and someone paid to look after it. The device had none of the three. So every practice that makes a service trustworthy, from being findable to being billable, grew up on the server and never left.
Devices carried none of those responsibilities for a simple reason: nobody called them. Your phone never needed to publish an interface because no other system called it. That absence was not negligence. It was the right answer to a question nobody had asked.
The question is being asked now. Agents are not chatbots that answer a single question. They are assistants given a goal, and they pursue it by calling tools, models and other agents, choosing the route as they go. When those agents live on devices, devices start serving requests from one another. The phone answers the car. The gateway on the factory floor answers the tablet in the supervisor’s hand. And the platforms that devices run today, whether the mobile operating system tuned for battery life or the container orchestration built for interchangeable racks of servers, were not designed for this world. Data center resources are generic; edge resources are typed. A camera is not swappable with another camera, and a phone is owned by a person, not by an operations team.
What mimOE actually does
mimOE is a small piece of software, tens of megabytes, that installs the server side of the connection on the device itself. It runs the same way on a phone, a gateway, a vehicle, a PC and a rack server. Once it is there, the device has a routable address, a discovery entry and a node identity, and it can serve requests from other nodes and from the agents that run on them.
Inside, it combines in one process what a data center assembles from a shelf of separate products: an API gateway that authenticates and authorizes each call, a microservice runtime, an AI inference runtime with an OpenAI-compatible API, and the discovery and observability that let nodes find and watch one another. We call those microservices mims. Each one runs in its own sandbox, JavaScript or WASM, on its own thread inside the mimOE process, and it is serverless and request-scoped: it starts for one request and nothing stays resident when it is done. A fault in one sandbox does not spread to the next, and a crash costs only the invocations in flight.
The agents themselves are mims. So are the higher functions of the platform, such as the ai-router that reads the context of a request and decides which inference should answer it, the loader and the mapreducer. Inference runs on the device on llama.cpp, ONNX or LiteRT, with the acceleration the hardware already provides, Metal, CUDA, Vulkan or ROCm, and with KV caching in the inference path. Tools reach an agent through an MCP gateway, where access is limited per agent by a security key. Images are defined in the OCI format and distribute from node to node, so a fleet updates itself without every device pulling from a central registry. When a device has no routable address, the tunnel gives it one; mimik operates the tunnel today, and a company can host its own.
The practical result is what we call the Device-First Continuum: intelligence starts on the device where the data is created and the decision is taken, and extends toward the cloud only when that is the better place for the work. An agent written once has a place to run everywhere along that continuum, including the places where no agent runs today.
The server is the means. The agent is the point. mimOE is an operating environment for agents, and every device that runs it becomes a place an agent can live.
The obligations arrive with the runtime
Here is the part that is easy to miss and impossible to skip. Installing the runtime does not remove the need to define interfaces, permissions, pricing and failure semantics. The moment a phone answers a request from another node, it needs a published interface, because a caller now depends on it. It needs accounting, because that caller consumed its CPU, its accelerator or its model. It needs to be upgradable while callers exist, and it needs defined failure semantics, because it will lose connectivity in the middle of a request.
Telecom operators wrote this checklist down decades before the web under the name FCAPS: fault management locates what broke, configuration management keeps settings and versions in order, accounting records usage per caller, performance management measures latency and capacity, security management controls access. The data center reinvented each item under newer names: observability, configuration management, chargeback, service-level objectives, identity and access management. The substance never changed: a component that others rely on owes them predictability.
mimOE carries the device-level part of that checklist natively. Isolation, discovery, identity, policy enforcement on every call, and attestation that the node is what it claims to be, all ship with the runtime. What remains is design work for the developer who writes the mim and the operator who runs the fleet. We are honest about that division because it is where our partners add their value.
The lesson from the first time we tried this
The web already ran a version of this experiment. Around 2005, when companies first opened programmatic access to their platforms, the promise was the mashup: combine existing services into something new. Property listings on a map. Weather on a calendar. A registry called ProgrammableWeb was founded to catalogue every available service, and the belief was that value would come from recombination rather than construction.
A mashup composed three or four APIs behind one interface: the user saw one convenient service, and several independent providers did the work. It ran into four problems that had nothing to do with HTTP.
Payment. When one service calls a second, which calls a third, and the customer pays only the first, nobody had a way to settle the chain. Everyone fell back on one-off contracts between pairs of companies, and those contracts do not scale.
Ownership of the combined offering. May the aggregator resell a provider’s output under its own name? Who keeps the user’s data afterward? Every combination was settled by terms of service, case by case, which made every combination legally fragile.
Management. There was no equivalent of FCAPS for API governance: no fault model across the chain, no accounting per caller, no end-to-end measure of performance. The API economy got rate limits and a status page, and a mashup that failed could not say which piece had failed.
Credit. The provider whose service did the work was invisible to the customer, so its only lever was advertising. Providers responded by closing their APIs and repricing access.
The registry outlived the revolution. ProgrammableWeb catalogued more than 19,000 services over seventeen years and closed in early 2023. Composition never became an industry.
Agents are mashups that assemble themselves
An agent is a mashup that composes itself: given a goal, it selects tools, models and other agents at runtime, with no human fixing the arrangement in advance. Everything the mashup era imagined assembling by hand, an agent performs at runtime, per request, and differently each time.
When every device is a server, that chain crosses owners by construction. The phone belongs to a person, the gateway to an operator, the workstation to an enterprise, the model to a vendor. Who pays whom when a request from an agent in a car is answered by a model on a phone and a cache on a gateway? Who owns the result when an agent, not a web page, presents it to the user? Who finds the fault when the pieces are mobile and intermittently connected? Who gets credit when the user sees only the agent?
The mashup era had hundreds of providers and could absorb the problem with bilateral deals for a while. The agent mesh will have millions of nodes and one new composition per request. What was a commercial nuisance for a few providers becomes a protocol problem for everyone. And it has to be solved at the device, because the device is where the call is served.
Two properties make a large system viable: traceability and non-anonymity. Every call must be attributable to an identified node, an identified agent and the user on whose behalf it was made, and every result must carry that provenance. The four problems above are one problem seen from four sides: payment, ownership, management and credit all require knowing who did what, for whom. Imagine a world where nothing is traceable and everyone is anonymous; people would line up to rob banks. A mesh of millions of anonymous nodes serving untraceable calls is that world, which is why identity and attestation sit at the node, in the runtime, and not in a policy document.
The answer cannot be the one from 2005: contracts, terms of service and marketing budgets. At the scale of the mesh, the four problems have to be solved in the protocol and enforced at the node.
Two problems, one runtime
mimOE is built to answer both problems this post has laid out, and it ships them together.
Placement. Agents do not yet run where the data is produced and the decision is taken. mimOE puts them there, as sandboxed, serverless mims with an on-device inference runtime behind an OpenAI-compatible API and the discovery to find one another across the mesh. An agent that runs on the node keeps the data on the node and answers without a round trip to a data center.
Discipline. The FCAPS checklist arrives on the device with the runtime. Identity for the node and for the agent. Association with the user and the account. Enforcement of policy on every call. Attestation. Fault isolation per mim. Software delivered as versioned images. Decisions that belong to the whole mesh, logical grouping, role definition, delegation and policy authoring, sit in servers above the node; the node enforces what they decide.
The table below puts the four old questions beside the answer the device now carries.
The old question | How the first API era answered it | How mimOE answers it at the device |
|---|---|---|
Who finds and fixes the failure? | Rate limits and a status page. | Each mim runs in its own sandbox, so a failure is contained and located. Software ships as versioned images. Every node reports its health and usage. |
Who pays whom? | Per-key quotas and one-off invoices between companies. | The device that does the work records each call and the resources it used (CPU, accelerator, model tokens), which is the record settlement between owners requires. |
Who is acting for whom? | Terms of service, negotiated case by case. | The node, the agent and the user are tied together, and the chain of delegation records on whose behalf each call was made. |
Who gets the credit? | Logo requirements written into contracts. | Provenance travels with the result: each response carries the verifiable identity of the node and the agent that produced it. |
Engage
The web had the interfaces and the registry, and it failed for lack of management, settlement and attribution. The agent mesh faces the same choice at a scale where the wrong answer is not recoverable, and the answer has to ship with the runtime that puts the agent on the device. That is what mimOE is for, and it is why we believe the next generation of AI companies will be built by the people who understand that capability and responsibility have to arrive together.
Captain Picard sent the Enterprise on every voyage with a single word. It applies here without modification. Put mimOE on the node, and let the agent go where no agent has gone before.


