Choose the model you want.
Run local models across Safetensors, MLX, and GGUF paths. Choose the model and execution path for the hardware and work at hand, from CPUs and GPUs to Apple Silicon.
The Inomina Gateway inference server runs local models on your machine, and the Inomina Mate agent system carries out multi-step computer work. Inomina brings the model, knowledge, workflow, and tools into one application.
QwenAlibaba
GemmaGoogle
ClaudeAnthropic
PLaMoPFN
From model choice and data location to network boundaries and real work, shape local LLMs around your work and environment.
Run local models across Safetensors, MLX, and GGUF paths. Choose the model and execution path for the hardware and work at hand, from CPUs and GPUs to Apple Silicon.
Work with conversations, ingested material, and finished output inside the environment you operate. Model and data stay together in the same environment.
Keep models and data on site and run inference without an external connection — documents, knowledge search, and operational workflows all complete inside the network. A perpetual license runs offline, so a closed network needs one; subscriptions re-verify online every 24 hours and stop once the grace period passes.
Combine them with internal knowledge, workflows, and the tools you already use—not just for chat, but for documents, search, research, and routine execution in one environment.
Do more than run a model. Put the local LLM you want to work in the way your job requires.
Putting a local model to work takes more than the model: an application to work in, a workspace for each kind of task, a knowledge layer, a workflow engine, an agent that operates a computer, and the inference service beneath them. Inomina implements all six and stacks them as a single design.
No layer is handed to somebody else's service. Underneath the inference layer we embed open-source execution engines, which Inomina builds and drives; the model lifecycle, the serving API, knowledge, workflows, and the agent above them are our own implementation. Because the boundaries between them belong to the same design, a model loaded once behaves identically in chat, in a workflow, and in the agent system — and changing that behaviour is a change to our own code, not a request to a supplier.
Use Gateway to discover, load, serve, and observe local models, while keeping cloud adapters available for workloads that need them.
Move between chat, translation, speech, knowledge search, documents, slides, reports, and research without rebuilding each experience from scratch.
Use the Mate agent system when a task needs planned browser, shell, file, search, or configured MCP actions—with approvals and execution evidence.
Inomina presents different working modes for different outcomes: everyday chat and media, translation and transcription, knowledge retrieval, document and presentation creation, deep research, and executable work through an agent system.
Select Standard, Image, Speech, Translation, Transcribe, Notebook, Global RAG, Agentic RAG, Document, Slide, Report, Deep Research, Fusion, or Mate according to the work in front of you. Availability can depend on configuration.
Bring in what the work already runs on — PDFs, Word, Excel, PowerPoint, plain text, a URL. Turn on OCR and scanned pages are read in Japanese and English; point it at a vision model and figures are described, so both become searchable text. Every answer carries citations, and each one opens the passage it came from.
A personal notebook and a shared collection are separate stores, not the same feature with a flag. Anyone signed in can read the shared side; creating collections, ingesting into them, and changing their settings stay with administrators — including through the API. Taking a collection out of the search set applies to everyone at once.
Notebook answers from the files you attached to this conversation. Global RAG answers from the collections your organization curates. Hybrid searches both at once, so a question that touches your draft and the company handbook does not have to be asked twice.
One search only answers a question you already knew how to phrase. Agentic RAG plans what to look for, searches, reads what came back, and rewrites its own queries where the evidence is thin — repeating until the plan is satisfied, the coverage gap closes, or nothing new turns up. The plan, each search, every replan, and the final answer stream into the conversation as they happen, so you can see why it concluded what it did.
Hierarchical summaries let a broad question match a summary while a specific one still matches the detail. An extracted entity graph pulls in material that shares no words with the question. Reranking reorders the shortlist. And the simulator shows what a question actually retrieves — with and without graph weighting, side by side — before you change anything for everyone.
Ingestion runs as a job with a visible step, so a bad result points at extraction, embedding, or storage rather than at "the AI". Documents can leave the search set without being deleted. Usage figures show which collections people actually rely on. Embedding and reranking always run on local models. Summarising and graph extraction use the retrieval model you select, so keeping that one local keeps indexing on your own hardware.
Build an executable graph from LLM, code, generated-tool, RAG, and MCP nodes. Connect inputs and outputs visually instead of hiding the whole process inside one prompt.
Add conditions, iteration, while loops, parallel branches, Try/Catch, delays, and merge nodes. Configure variables and ports, then inspect execution results from the same workflow.
Manage models, connect repeatable steps, and bring the tools your work already depends on.
Gateway brings model discovery, download, compatibility checks, loading, unloading, and service status into one control surface.
Start with manual input or a webhook, then connect LLM, RAG, and tool nodes with conditions, loops, parallel branches, retries, delays, and merges.
Call tools from configured MCP servers, use built-in code tools, or create a visual tool and place it directly in the workflow.
The application brings model selection, focused AI workspaces, knowledge, and visual workflows together. Gateway supplies inference; the Mate agent system supplies tool execution.
Workspace: chat, media, translation, knowledge, documents, research, and workflows.
Gateway: model routing, local runtimes, model lifecycle, and inference APIs.
Agent System: planned browser, shell, file, search, and MCP execution with approvals and observable output.
01 / MODEL LAYER
Gateway
local models · cloud adapters · inference
02 / WORKSPACE
Inomina
chat · RAG · documents · visual workflows
03 / EXECUTION
Agent System
browser · shell · files · MCP
Inomina is built on two dedicated systems with clearly separated roles: Inomina Gateway runs the models, and Inomina Mate executes the work.
The inference server that runs local LLMs. It executes GGUF and Safetensors models on your CPU or GPU, exposes them through app-ready APIs, and manages everything from discovery and download to loading and release.
Learn More →The agent system that does computer work on your behalf. It breaks a goal into a plan and executes it with the browser, shell, files, search, and configured MCP tools — progress, approvals, artifacts, previews, and diffs are visible per session.
Learn More →Select and operate a local model through Gateway, with cloud access remaining an explicit option rather than the premise.
Use that model in chat, translation, knowledge retrieval, document creation, research, or a visual workflow.
When the task requires action, hand it to the Mate agent system with a session, execution policy, approvals, and observable outputs.
The way you work together changes; the application does not. Only what it connects to changes with you.
Everything runs on your own computer. The model, the conversations, and the documents stay where you put them.
Personal license · single seatServe one machine on your network and let colleagues share the same models and knowledge from their own desktops.
Team license · five seatsSeats and concurrent sessions follow your agreement, and members can sign in from a browser without installing the desktop app.
Enterprise license · contract seatsCompare what each license includes on the pricing page.
Adopt one capability first, then connect the others when the work requires them.
Begin by serving one local model for a concrete chat, extraction, coding, retrieval, translation, or speech task.
Use the model in a focused workspace or connect model, knowledge, tools, and control flow visually.
Add the Mate agent system when the desired result requires planned work outside a model response.
Choose where the model should run, select the workspace for the outcome, and add a workflow or the Mate agent system only when the task needs execution.
Get Started