palOMine · Personal agentic AI appliance
Your own AI. Your own hardware. Your agents work for you.
palOMine is a local-first agentic AI platform packaged as a personal appliance.
Reasoning, coding, memory, voice, image generation, retrieval, and background agents all run on hardware you own. No subscription API required. Your prompts, memory, embeddings, and agent state stay on your network.
And palOMine is not built around making one giant model do everything.
Join the waitlist for the first production run.
Compute lanes
One box. Multiple inference lanes.
An agentic system should not have to stop everything whenever one model is thinking. palOMine uses the hardware as a set of specialized compute lanes.
Primary reasoning
NVIDIA Nemotron 3.5 Lightning 30B-A3B
runs on the integrated GPU for reasoning, coding, planning, and complex agent work.
Background intelligence
Qwen3.5-2B
runs independently on the NPU for lightweight classification, routing, extraction, summarization, and other background tasks.
Supporting models
Whisper · Kokoro · FLUX.2 · RealESRGAN · nomic-embed · bge-reranker
speech recognition, speech synthesis, embeddings, reranking, image generation, and upscaling distribute across the CPU and integrated GPU as their own lanes.
That means the primary reasoning model does not need to spend its time doing every small piece of work itself.
One appliance
GMKtec EVO X-2, 128 GB.
One appliance. One supported configuration. 128 GB of unified memory gives palOMine enough room to keep its primary model, agent state, memory system, and supporting AI services local.
No separate GPU to install. No rack of accelerators. No proprietary cloud runtime. Plug it in, connect the services you want to use, and the platform is ready.
Included AI stack
- Reasoning & coding
- NVIDIA Nemotron 3.5 Lightning 30B-A3B
- NPU background model
- Qwen3.5-2B
- Speech to text
- Whisper Large v3 TurboMoonshine Medium Streaming
- Text to speech
- Kokoro
- Image generation
- FLUX.2 Klein
- Image upscaling
- RealESRGAN x4+
- Embeddings
- nomic-embed v1
- Reranking
- bge-reranker v2 m3
These are part of the appliance configuration, not a catalogue of models customers are expected to configure themselves.
Real-world performance
Real-world agent performance.
~57 tok/s sustained through a long-running agent workload.
Not a one-prompt speed test.
In our latest agent run, Nemotron completed 36 inference turns and generated more than 36,000 output tokens while the active working context grew beyond 40,000 tokens.
Across substantive generations, the model sustained approximately 57 tok/s.
One turn generated 9,770 tokens at 57.26 tok/s. Another generated 4,682 tokens at 56.46 tok/s.
The point is not the peak.
The point is that it keeps doing it.
- Sustained generation
- ~57.0 tok/s
- TTFT (cached)
- ~0.58s median across the agent run
- Background NPU inference
- Qwen3.5-2B · ~25 tok/s
- Workload model
- Nemotron 3.5 Lightning 30B-A3B MXFP4
Built around
Why this matters
A useful local AI system is not defined by the fastest single benchmark — it needs to stay responsive after hours of work, memory that does not grow forever, tools and agents that can operate without handing every task to the largest model, and enough independent compute to do more than one kind of work.
That is what palOMine is being built around.
One appliance. Multiple models. Multiple compute lanes. One agent platform.
Local-first
Your data stays here.
- · Prompts stay local.
- · Embeddings stay local.
- · Agent memory stays local.
- · Run state stays local.
- · Monitoring stays local.
The appliance can connect to external services when you explicitly ask it to, but the AI platform itself does not require your conversations to pass through somebody else’s inference API.
Reserve a unit for the first production run.
palOMine is currently preparing for its first hardware run. Join the waitlist and we’ll notify you before units become generally available.