Skip to content

Benchmarks

  • +109%

    More KV cache (2.09x). Gemma 4 12B (BF16) on RTX 5090.

  • +39%

    More KV cache (1.39x). Qwen 3.8-27B (BF16) on RTX PRO 6000 Blackwell.

  • Exact

    Bit-exact weights without quantization. Exact model output.

Benchmarks on GitHub

Models

A few models with varying architectures and sizes. Smaller .tic leaves more memory for KV cache. BF16 is supported today, with more precisions in progress.

Compile your own modelsModels on Hugging Face

ModelPrecisionTypeRaw.tic (savings)
Qwen2.5 7B Instruct on Hugging FaceBF16Text15.23 GB10.86 GB (28.7%)
Mistral 7B Instruct v0.3 on Hugging FaceBF16Text14.50 GB10.33 GB (28.8%)
Llama 3.1 8B Instruct on Hugging FaceBF16Text16.06 GB11.44 GB (28.8%)
Gemma 4 12B IT on Hugging FaceBF16Multimodal23.92 GB17.05 GB (28.7%)
Qwen3.5 27B on Hugging FaceBF16Multimodal55.56 GB39.57 GB (28.8%)
Qwen3.5 35B A3B on Hugging FaceBF16Multimodal, MoE71.90 GB51.27 GB (28.7%)
Qwen3.8 27B on Hugging FaceBF16Multimodal55.56 GB39.57 GB (28.8%)