Image Distillation
Your house style, at one to four steps and a fraction of the cost — a model you own, not an API you rent. Distilled on your catalogue, verified by our judge model, deployed your way.
Your house style in one to four steps
Guidance is where the money goes. The teacher runs the full schedule and pays twice at every step to hold a prompt; a student distilled on your catalogue and your brand adapters folds that guidance into its weights and lands the same image in one to four steps. The product shots and campaign assets you generate by the thousand cost a fraction of that, and still look like you.
The architecture
Six stages, from your own data to a model that belongs to you — and a loop that keeps it current as your traffic moves.
-
01
From your catalogue to training pairs
We ingest product shots, campaign assets and past generations with the prompts that made them, drop near-duplicates by perceptual hash and embedding distance, cluster on CLIP embeddings to see what the library covers, recaption with a VLM, and hold out a stratified slice per cluster.
perceptual-hash dedupeCLIP embedding clustersVLM recaptioningaesthetic pre-filterheld-out eval slice -
02
The expensive pass runs once
Your teacher — SDXL, SD3 or FLUX-class, or your own fine-tune — samples the full schedule under guidance while we harvest solver trajectories, noise–latent pairs for reflow and CFG-combined predictions. Latents and embeddings cache beside them; the frozen teacher stays loaded for the online terms.
CFG-combined targetssolver trajectoriesnoise–latent pairscached VAE latentsfrozen teacher score -
03
Four ways to score the same image
Consistency alone softens texture; an adversarial term alone narrows diversity. So they run together — latent consistency for the step collapse, distribution matching between teacher and learned student scores, an adversarial head for detail, reflow to straighten paths, guidance distilled in.
LCM consistency lossdistribution matchingadversarial (ADD/LADD)rectified-flow reflowguidance distillation -
04
Small enough to run at volume
Depth and width come down, attention runs in fused kernels, weights land at int8 — fp8 where the hardware has it — and the text encoder is quantised too. A distilled VAE decoder keeps pixels from becoming the new bottleneck. Brand LoRAs are refit on the student, merged in or swappable per line.
block pruningint8 / fp8 weightsdistilled VAE decoderbrand + style LoRAsswappable adapters -
05
The bar an image has to clear
Every candidate runs a prompt suite from your briefs and the held-out slice. Our judge model scores it against teacher output on adherence, brand, aesthetics and diversity; a regression set re-runs past failures; likeness, trademark and NSFW filters run on every sample. One gate red, nothing ships.
Zorbe judge modelprompt-adherence evaldiversity + aestheticsbrand-conformance setsafety + NSFW filters -
06
Where it runs, and what returns
It ships hosted or into your own VPC behind a versioned endpoint — weights pinned, rollback one call away, overnight catalogue batches and interactive requests on the same build. The prompts you send, the variants your team picked and the ones an editor killed flow back into the corpus.
hosted or in-VPCpinned versionsone-call rollbackpicked / killed pairscorpus refresh
aYour domain
Distilled on your own catalogue — your products, your palette, your house style — so it is right on your assets, not on everything.
bJudge-verified
The student ships only when our judge model confirms it matches the teacher on adherence, brand and aesthetics, behind canary gates.
cDeployed your way
Hosted by us or inside your own environment — a fraction of the cost and latency at catalogue volume, and the model is yours.
Zorbe