palOMine surfaces

An image pipeline that lives on the box.

Image generation on palOMine runs the same way the rest of the appliance runs — locally, on the hardware, with no third-party image API in the loop. FLUX.2 Klein handles text-to-image. RealESRGAN x4+ handles upscaling. Both render on the appliance under the same bounded-authority agent as every other surface.

The local image stack

Two models. No cloud image API.

Every image generated or refined on palOMine passes through one of two on-box models. None of them call a hosted image generation service. The full list of preloaded models lives on /models.

  • Text-to-image

    FLUX.2 Klein

    Generates images from natural-language prompts. The agent calls it when a turn warrants a visual — concept sketches, illustrations for a draft, figures to accompany a chat reply — and the rendered bytes stay on the appliance.

    A small, fast open-weight text-to-image model selected to fit the appliance memory budget alongside the rest of the model set.

  • Upscaling

    RealESRGAN x4+

    Boosts a generated or supplied image up to 4× its linear resolution without re-running the diffusion path. The on-demand refinement step for output that needs to be sharper than the initial render.

    Single-pass super-resolution — practical for clean-up after generation or for rendering a small output at presentation size, not a substitute for the diffusion model itself.

Typical usage

Prompt in, image out. On demand, on box.

The image stack is not a separate product. It is a tool the agent calls when a turn warrants a visual. Three flows cover most of what it gets used for.

  • Prompt-driven generation

    A user asks the agent for an image inline — a sketch of a part, an illustration next to a project note, a draft hero for a deck. The agent forwards the prompt to FLUX.2 Klein on the appliance; the rendered pixels never travel further than the local conversation store.

  • On-demand upscaling

    Once an image is generated — or one supplied to the agent — RealESRGAN x4+ refinements it on demand when a higher-resolution version is wanted for embedding, export, or display. It is a single, fast pass, not a re-render through diffusion.

  • Agent-mediated, appliance-controlled

    Image generation and upscaling are tool calls on the same bounded-authority agent that handles typed turns. Whether the user typed the prompt, spoke it via voice, or sent it through a chat gateway, the appliance decides when the tool is allowed to run.

Composes with the chat agents + bounded-authority core

Agents can generate images when authorised. Outputs stay local.

Image generation is a tool call on the same agent that handles a typed turn. The chat gateways — /features/channels — and the on-device web UI can both ask the agent for an image inline; the agent decides whether the call is permitted under the appliance’s allow/deny rules before FLUX.2 Klein runs.

The rendered pixels never leave the LAN. The appliance controls what the agent is allowed to generate, where the output lands, and whether it can be sent onward through a chat gateway. The image stack is a tool, not an autonomous surface — same model as every other tool call the agent can make.

Privacy

Prompts and outputs never leave the box.

There is no cloud image generation API. There is no cloud upscaling API. Text-to-image inference runs on FLUX.2 Klein on the appliance; the refinement pass runs on RealESRGAN x4+ on the appliance. Prompts go in; rendered pixels come out; nothing egresses to a third-party image service in either direction.

Generated images and their prompts live on the appliance, alongside the rest of the conversation archive. If image generation is disabled on a surface or on the device as a whole, the tool call stops with it. Nothing the appliance renders is replayed through a hosted model.

Capability expectations

What the appliance actually delivers.

FLUX.2 Klein is a small, fast text-to-image model selected to fit the appliance memory budget alongside the rest of the model set. It is not a hosted-grade photoreal engine, and it is not positioned as one. The upscaler delivers a single-pass 4× refinement on top of that initial render — sharper output at presentation size, without re-running diffusion.

Practical expectations: prompts in the normal conversational length range, image sizes sized to the appliance’s working memory, time-to-first-pixel in seconds rather than the latency of a hosted model. The goal is on-box availability and privacy, not parity with a cloud image service. The full preloaded set is listed on /models.

The image stack ships with the first run.

FLUX.2 Klein and RealESRGAN x4+ are part of the production model set on every appliance. Get on the waitlist and we’ll let you know when units are ready to be reserved.