Developer reference

The models palOMine ships with.

Nine open-weight models, pre-tuned end-to-end on a single GMKtec EVO X-2. No API keys, no cloud — every weight below runs on the appliance, in your network, behind your gate.

What this page is

A reference doc — not a marketplace.

These are the models the appliance ships with. They can be changed or swapped, but doing so is unsupported and untested — palOMine is pre-tuned end-to-end on this stack, so compatibility and performance outside it are not guaranteed. This page documents the actual stack for developers evaluating the device — what each model is, what the appliance uses it for, and what it is good at.

  • Main reasoning & coding agent

    NVIDIA Nemotron-3.5-Lightning (30B, quantized GGUF)

    The primary agent that reads, edits, and runs code. Plans multi-step work, drives the coding assistant, and answers long-form reasoning prompts.

    Capability: Large-context code + reasoning; the heaviest model on the box.

  • Background / secondary model

    Qwen3.5-2B

    Handles lightweight side-tasks — short classifications, summaries, watchdogs — so the main agent stays focused on the user turn.

    Capability: Small, cheap, and fast — the appliance default for anything that is not the main agent.

  • Speech-to-text

    Whisper-Large-v3-Turbo

    Transcribes dictated prompts and chat messages from the chat gateways and the web UI. First stage of every voice-in surface.

    Capability: Multilingual ASR with a good latency/quality trade-off.

  • Text-to-speech

    Kokoro

    Speaks agent responses back through the chat gateways and the on-device web UI. The voice half of every voice-out surface.

    Capability: Low-latency TTS, sized for natural conversational pacing.

  • Streaming speech recognition

    Moonshine-Medium-Streaming

    Low-latency live captioning-style input — listens while the user speaks instead of waiting for a finished utterance, for dictation flows that need to feel reactive.

    Capability: Streaming STT — incremental, turn-taking-friendly transcription.

  • Image generation

    FLUX-2-Klein

    Generates images from prompts inside the appliance — illustrations, diagrams, generated artwork — entirely on-device.

    Capability: Small, fast image generation tuned for the appliance memory budget.

  • Image upscaling

    RealESRGAN-x4plus

    Upscales generated or imported images 4× for higher-resolution outputs without re-running the diffusion model end-to-end.

    Capability: Super-resolution — sharpens low-resolution sources before they leave the box.

  • Embeddings

    nomic-embed-text-v1

    The vector backbone for retrieval and long-term memory: every note, file, and conversation gets projected into a vector store on the appliance.

    Capability: Text embeddings, sized for high-quality on-device retrieval.

  • Reranking

    bge-reranker-v2-m3

    Reranks candidate passages returned by the embedding-based retriever so the top results the agent sees are actually relevant, not just nearest-neighbour.

    Capability: Cross-encoder reranker — the precision step after vector recall.