# Konic > Konic builds production-optimized LLMs on infrastructure you own. Three model families — Uno-1 (open source), Duo-1 (enterprise production), Tres-1 (mission-critical) — plus custom LLM development, delivered under an annual licence with no per-token cost, on-premise or private cloud. Last updated 2026-09-08. Konic builds compact, production-optimized LLM families for enterprises that need AI inside their own security boundary. Instead of per-token API pricing on models sized for every possible task, Konic engineers models for the workload you actually run — pruning, distillation, and quantization targeting your production hardware — and licenses them annually on machines you control. Open-source research on model compression (expert pruning, AWQ INT4 quantization, distillation alignment) backs every claim; published results include ~91% performance retained after optimization, 4x smaller models for the same task, and 6.1x memory reduction. ## Model families - [Konic Uno-1](https://konic.io/models/uno-1): Open-weight model family published for developer adoption. Free to download, self-host, and evaluate on your own hardware. - [Konic Duo-1](https://konic.io/models/duo-1): Enterprise production family. Balanced capability for production workloads, delivered under an annual enterprise licence on on-prem infrastructure you own — no per-token cost. Commercial terms scoped per engagement. - [Konic Tres-1](https://konic.io/models/tres-1): Mission-critical family. Highest capability tier where accuracy carries the most risk — banking, insurance, healthcare, manufacturing. Annual enterprise licence, air-gap capable, managed inference optional. ## Deployment Deployment paths: self-hosted in your own environment (private networking, your keys, your data plane), private cloud VPC, edge and constrained hardware (pruning and quantization target real production hardware including Apple Silicon and GPUs), air-gapped for regulated and sovereign workloads. Versioned family releases keep integrations stable across upgrades. Managed inference available for teams that prefer Konic to run serving. ## How Konic works Working model: choose one AI flow already running in production, deploy a Konic model family inside your environment, A/B test against the incumbent model or your success criteria, decide on measured results. Custom LLM development builds task-specific models from your data and requirements, delivered to infrastructure you control. ## Research Published engineering with reproducible results: - [Optimized LFM2.5-VL-3B: FFN Width Pruned, Distillation Aligned, INT4 Quantized](https://konic.io/research/lfm2.5-vl-3b-optimized): 2026-08-13 - Joint-SwiGLU pruning and INT4 quantization of a 3B vision-language model; live comparison preserving tool calling. - [LFM2.5 Encoder 230M + SigLIP2: A Compact Multimodal Encoder](https://konic.io/research/lfm2.5-multimodal-encoder-230m): 2026-08-12 - 230M-parameter encoder-only multimodal model with GPTQ INT4 release at -59.9% package size. - [Two Stages, Much Smaller MoE: REAP Expert Pruning Followed by AWQ INT4 Quantization](https://konic.io/research/reap-awq-moe-compression): 2026-07-20 - 32→16 expert pruning plus AWQ INT4 on a MoE model; 2.79 GB packed artifact. - [From-Scratch AWQ INT4 Quantization on Qwen3-8B](https://konic.io/research/awq-int4-qwen3-8b): 2026-07-03 - 4.0x smaller model at 1.034x FP16 perplexity; norm-folding the AWQ scale is the decisive factor. - [REAP Expert Pruning for On-Device MoE Models](https://konic.io/research/reap-expert-pruning-apple-silicon): 2026-06-29 - Router-weighted expert pruning on Apple Silicon; 96.8% code-gen performance at 25% compression. ## Guides Educational content on on-prem LLM deployment: - [LLM Data Sovereignty: Why Enterprises Keep Models In-House](https://konic.io/blog/llm-data-sovereignty): 2026-09-08 - Why data sovereignty drives on-prem adoption and how to evaluate a deployment against it. - [Air-Gapped LLM Deployment: What It Means and When You Need It](https://konic.io/blog/air-gapped-llm): 2026-09-05 - What air-gapped deployment requires and which industries need it. - [How to Deploy an LLM On-Premise: A Practical Guide](https://konic.io/blog/deploy-llm-on-premise): 2026-09-03 - Model choice, hardware sizing, serving runtimes, and quantization. - [On-Prem LLM vs API: When to Self-Host Your Language Models](https://konic.io/blog/on-prem-llm-vs-api): 2026-09-01 - Cost, latency, privacy, and operational trade-offs. ## Statistics | Metric | Value | Source | Date | |---|---|---|---| | 91% perf retained | REAP 25% comp | https://konic.io/research/reap-expert-pruning-apple-silicon | 2026-06-29 | | 4.0x smaller 1.034x PPL | AWQ Qwen3-8B | https://konic.io/research/awq-int4-qwen3-8b | 2026-07-03 | | 2.79GB MoE artifact | REAP+AWQ | https://konic.io/research/reap-awq-moe-compression | 2026-07-20 | | -59.9% package | encoder GPTQ INT4 | https://konic.io/research/lfm2.5-multimodal-encoder-230m | 2026-08-12 | | -69% 1.93GB tool-calling kept | LFM2.5-VL-3B | https://konic.io/research/lfm2.5-vl-3b-optimized | 2026-08-13 | ## Pricing and licensing terms - Uno-1: open weights, free to download and self-host from https://huggingface.co/konic-labs — no licence fee. - Duo-1 and Tres-1: annual enterprise licence per model family, flat fee per host, no per-token cost; versioned releases with support tiers; deployed on infrastructure you own. - Commercial terms including DPA and SLA are scoped per engagement (request access via https://konic.io/pricing). - Custom LLM development: scoped per engagement (task definition, data preparation, training/adaptation, evaluation, serving integration, release planning). - Pricing page: https://konic.io/pricing ## Company - [Pricing](https://konic.io/pricing): Annual licence per model family; custom model development scoped per engagement. - [Custom LLM development](https://konic.io/solutions/custom): Domain-specialized models built from company data for agentic pipelines. - [About Konic](https://konic.io/): NVIDIA Inception Program member. Founded by Gokalp Katkat (Co-founder, CEO, gokalp@konic.io) and Ege Sabanci (Co-founder); pre-seed stage. ## Links - [GitHub — konic-labs](https://github.com/konic-labs) - [Hugging Face — konic-labs](https://huggingface.co/konic-labs) - [LinkedIn — Konic](https://www.linkedin.com/company/koniclabs) - [X — @koniclabs](https://x.com/koniclabs) - [Book a meeting](https://calendar.app.google/BA7ADCjsyesnHL1Z8) ## Contact - Gokalp Katkat (Co-founder, CEO) — gokalp@konic.io - Ege Sabanci (Co-founder) - General — via https://konic.io or book a meeting above.