• Konic Uno-1Open-weight family — download, self-host, evaluate.
  • Konic Duo-1Enterprise production family — annual enterprise licence, on-prem.
  • Konic Tres-1Mission-critical tier, highest accuracy bar.
  • All model familiesCompare Uno-1, Duo-1, and Tres-1 — deployment and licence options.
  • Custom LLM DevelopmentLLMs built for your agentic pipelines from company data and requirements.
  • Open Models on HFExplore open-weight releases on Hugging Face.
  • ResearchBenchmarks, compression notes, and release write-ups from the lab.
  • Optimized LFM2.5-VL-3B: FFN Width Pruned, Distillation Aligned, INT4 QuantizedWe surgically compress Liquid AI’s LFM2.5-VL-3B — an architecture-searched hybrid conv+attention vision-language model — by cutting FFN width where the search never optimized it: joint-SwiGLU Wanda pruning (10752→8192 and →7168), distillation-to-baseline LoRA recovery, and hand-rolled W4A16 GPTQ + W8 mixed-precision quantization. All three tiers are near-lossless against the base on our paired eval shards — PPL ratios at or below baseline (0.79–0.85), top-5 token agreement ≥ 0.95, quantization adding only +0.02–0.06 nats — across 5.30 GB (−15%), 4.92 GB (−21%), and 1.93 GB packed / 2.22 GB (Konic Optimized) — native INT4/INT8 compressed-tensors in vLLM (−69%). Live comparison on the smallest tier preserves scenes, facts, and tool calling — 10/10 tool rounds, including correct abstention where the base over-triggered. Recorded runs, not general benchmarks.
  • LFM2.5 Encoder 230M + SigLIP2: A Compact Multimodal EncoderWe augment Liquid AI’s LFM2.5 Encoder 230M — a 230M-parameter bidirectional masked-language encoder — with a SigLIP2 vision tower and 32 learned soft tokens to build an encoder-only multimodal model for retrieval and image-text matching. Clean image-only retrieval on 12,500 held-out MONET pairs reaches 0.1194 image→text R@1 with the BF16 reference; the GPTQ INT4 release retains 0.1091 while shrinking the package from 923.65 MB to 370.46 MB (−59.9%). Cyclic-negative matching AUROC is 0.9733; text-nearest hard negatives and image-conditioned masked-token prediction remain at chance, and the model does not generate.
  • Two Stages, Much Smaller MoE: REAP Expert Pruning Followed by AWQ INT4 QuantizationWe chain two complementary compression stages on Liquid LFM2.5-8B-A1B — REAP CUDA expert pruning (32→16 experts across 22 MoE layers) then external AWQ INT4 quantization (W4A16_ASYM, group 128). The published packed artifact is 1.15B packed-weight equivalent / 2.79 GB; recorded MATH500 and BFCLv3 results use a separate 9.18 GB AWQ-scaled BF16 derivative for vLLM evaluation, not direct packed-INT4 runtime. These are recorded run summaries, not general benchmarks.
  • From-Scratch AWQ INT4 Quantization on Qwen3-8BWe validate a from-scratch, pure-PyTorch AWQ implementation on Qwen3-8B: group-wise INT4 with per-channel AWQ scaling produces a 4.0× smaller model (13.9 GB → 3.5 GB linear weights) at 1.034× FP16 perplexity on WikiText-2 (10.08 vs 9.75), loaded and run in a real INT4 GEMM runtime. The decisive factor is norm-folding the AWQ scale — 20× more accurate per weight than weight-dequantization.
  • On-Prem LLM GuidesDeployment, cost, data sovereignty, and air-gapped LLM guides.
  • LLM Data Sovereignty: Why Enterprises Keep Models In-HouseData sovereignty is the reason most regulated enterprises cannot use hosted LLM APIs: the moment proprietary data crosses a network boundary, control over it is shared. This guide explains what data sovereignty means for AI, why it is driving on-prem adoption, and how to evaluate a deployment against your sovereignty requirements.
  • Air-Gapped LLM Deployment: What It Means and When You Need ItAn air-gapped LLM runs on infrastructure with no connection to the public internet — the model, the data, and the serving stack all live inside your boundary. This guide explains what air-gapped deployment actually requires, which industries need it, and the practical constraints of running a model with no external dependencies.
  • How to Deploy an LLM On-Premise: A Practical GuideDeploying an LLM on your own infrastructure is a sequence of concrete decisions: pick the model, size the hardware, choose a serving runtime, quantize for your GPU, and wire in monitoring. This guide walks each step with the trade-offs that actually matter for a production on-prem deployment.
  • On-Prem LLM vs API: When to Self-Host Your Language ModelsAPI-based LLMs are fast to adopt but scale poorly on cost, latency, and data control. On-prem LLMs trade setup effort for predictable economics, lower per-token cost at volume, and data that never leaves your boundary. This guide breaks down the decision across cost, latency, privacy, compliance, and operational effort — and when each path is the right call.
  • Book a demoTalk with the team about your workload.
  • GitHubOpen research and tooling.
  • Hugging FaceModels and model cards.
  • XFollow @koniclabs on X.
Pricing
Models on HFGitHubX
Book a demo
Book a demo

About Konic

Production-optimized LLMs on infrastructure you own.

Konic Labs builds compact, production-optimized LLM families for enterprises that need AI inside their own security boundary. Instead of per-token pricing on models sized for every possible task, we engineer models for the workload you actually run — pruning, distillation, and quantization targeted at your production hardware — and license them annually on machines you control.

Book a demoSee the models

Why Konic exists.

Enterprises rarely stall because a model is not good enough. They stall on what it costs and takes to run in production. API dependency means cost scaling with usage and data leaving the business. Raw open weights mean the buyer owns the compression, post-training, and serving engineering.

Konic removes that middle layer. Model families are engineered for the workload the buyer actually runs — sized for the task, not the benchmark — delivered as versioned releases, and licensed annually on the customer's own machines: no per-token cost, no usage-scaling bill, no data egress.

Every claim is backed by published engineering with reproducible results. Read the research or see how we build.

Published results

performance retained after optimization
~91%
smaller model for the same task
4x
memory reduction in published work
6.1x
founders, pre-seed, NVIDIA Inception
2

Team

Two founders, one compression pipeline.

Co-founder, CEO

Gokalp Katkat

Gokalp works across the model compression pipeline — domain adaptation, pruning and distillation, and the INT4 artifacts that ship for on-prem serving.

[email protected]

Co-founder

Ege Sabanci

Ege works across the model compression pipeline — domain adaptation, pruning and distillation, and the INT4 artifacts that ship for on-prem serving.

Company facts.

NVIDIA Inception Program member

Konic is an NVIDIA Inception Program member.

Legal name
Konic Labs, Inc.
Headquarters
Istanbul, TR
Focus
On-prem LLM deployment · Enterprise LLM · LLM compression · Model quantization · Custom LLM development · Sovereign AI
Contact
[email protected]

Elsewhere

  • GitHub
  • Hugging Face
  • LinkedIn
  • X
  • llms.txt
Konic

Compact production-optimized LLMs on infrastructure you own.

NVIDIA Inception Program memberAnnouncement (LinkedIn)

Models

  • Konic Models
  • Custom LLM Development
  • Pricing
  • Open Models

Research

  • How We Build
  • About Konic
  • Glossary
  • Research
  • Optimized LFM2.5-VL-3B: FFN Width Pruned, Distillation Aligned, INT4 Quantized
  • LFM2.5 Encoder 230M + SigLIP2: A Compact Multimodal Encoder
  • Two Stages, Much Smaller MoE: REAP Expert Pruning Followed by AWQ INT4 Quantization

Guides

  • On-Prem LLM Guides
  • LLM Data Sovereignty: Why Enterprises Keep Models In-House
  • Air-Gapped LLM Deployment: What It Means and When You Need It
  • How to Deploy an LLM On-Premise: A Practical Guide

Connect

  • Book a demo