Sample models
A few sample models with varying architectures and sizes. Smaller .tic leaves more memory for KV cache. BF16 is supported today, with more precisions in progress.
| Model | Precision | Type | Raw | .tic (savings) |
|---|---|---|---|---|
| Qwen2.5 7B Instruct on Hugging Face | BF16 | Text | 15.23 GB | 10.86 GB (28.7%) |
| Mistral 7B Instruct v0.3 on Hugging Face | BF16 | Text | 14.50 GB | 10.33 GB (28.8%) |
| Llama 3.1 8B Instruct on Hugging Face | BF16 | Text | 16.06 GB | 11.44 GB (28.8%) |
| Gemma 4 12B IT on Hugging Face | BF16 | Multimodal | 23.92 GB | 17.05 GB (28.7%) |
| Qwen3.5 27B on Hugging Face | BF16 | Multimodal | 55.56 GB | 39.57 GB (28.8%) |
| Qwen3.5 35B A3B on Hugging Face | BF16 | Multimodal, MoE | 71.90 GB | 51.27 GB (28.7%) |
Next steps
Each compiled model ships with a hash manifest for verification.
Benchmarks include memory, latency, and throughput.