palOMine · Personal agentic AI appliance

Your own AI. Your own hardware. Your agents work for you.

palOMine is a local-first agentic AI platform packaged as a personal appliance.

Reasoning, coding, memory, voice, image generation, retrieval, and background agents all run on hardware you own. No subscription API required. Your prompts, memory, embeddings, and agent state stay on your network.

And palOMine is not built around making one giant model do everything.

Join the waitlist for the first production run.

Compute lanes

One box. Multiple inference lanes.

An agentic system should not have to stop everything whenever one model is thinking. palOMine uses the hardware as a set of specialized compute lanes.

  • Primary reasoning

    NVIDIA Nemotron 3.5 Lightning 30B-A3B

    runs on the integrated GPU for reasoning, coding, planning, and complex agent work.

  • Background intelligence

    Qwen3.5-2B

    runs independently on the NPU for lightweight classification, routing, extraction, summarization, and other background tasks.

  • Supporting models

    Whisper · Kokoro · FLUX.2 · RealESRGAN · nomic-embed · bge-reranker

    speech recognition, speech synthesis, embeddings, reranking, image generation, and upscaling distribute across the CPU and integrated GPU as their own lanes.

That means the primary reasoning model does not need to spend its time doing every small piece of work itself.

One appliance

GMKtec EVO X-2, 128 GB.

One appliance. One supported configuration. 128 GB of unified memory gives palOMine enough room to keep its primary model, agent state, memory system, and supporting AI services local.

No separate GPU to install. No rack of accelerators. No proprietary cloud runtime. Plug it in, connect the services you want to use, and the platform is ready.

Included AI stack

Reasoning & coding
NVIDIA Nemotron 3.5 Lightning 30B-A3B
NPU background model
Qwen3.5-2B
Speech to text
Whisper Large v3 TurboMoonshine Medium Streaming
Text to speech
Kokoro
Image generation
FLUX.2 Klein
Image upscaling
RealESRGAN x4+
Embeddings
nomic-embed v1
Reranking
bge-reranker v2 m3

These are part of the appliance configuration, not a catalogue of models customers are expected to configure themselves.

palOMineEVO X-2 · 128GB
online · 128GB unified memory
agent · skills · memory · USB-C · USB-A · HDMI · RJ45 · SD · PWR

Real-world performance

Real-world agent performance.

~57 tok/s sustained through a long-running agent workload.

Not a one-prompt speed test.

In our latest agent run, Nemotron completed 36 inference turns and generated more than 36,000 output tokens while the active working context grew beyond 40,000 tokens.

Across substantive generations, the model sustained approximately 57 tok/s.

One turn generated 9,770 tokens at 57.26 tok/s. Another generated 4,682 tokens at 56.46 tok/s.

The point is not the peak.

The point is that it keeps doing it.

Sustained generation
~57.0 tok/s
TTFT (cached)
~0.58s median across the agent run
Background NPU inference
Qwen3.5-2B · ~25 tok/s
Workload model
Nemotron 3.5 Lightning 30B-A3B MXFP4

Built around

Why this matters

A useful local AI system is not defined by the fastest single benchmark — it needs to stay responsive after hours of work, memory that does not grow forever, tools and agents that can operate without handing every task to the largest model, and enough independent compute to do more than one kind of work.

That is what palOMine is being built around.

One appliance. Multiple models. Multiple compute lanes. One agent platform.

Local-first

Your data stays here.

  • · Prompts stay local.
  • · Embeddings stay local.
  • · Agent memory stays local.
  • · Run state stays local.
  • · Monitoring stays local.

The appliance can connect to external services when you explicitly ask it to, but the AI platform itself does not require your conversations to pass through somebody else’s inference API.

Reserve a unit for the first production run.

palOMine is currently preparing for its first hardware run. Join the waitlist and we’ll notify you before units become generally available.