Overview
Efficiency without loss
Compile models once into .tic with smaller memory footprint, then deploy through ISIRO Runtime, which integrates the inference frameworks you already use.
Supports BF16 vLLM on NVIDIA GPUs today. More precisions, frameworks, and hardware are in progress.
Measured results
Up to 30%
Smaller weights, more KV cache
Exact
Bit-exact weights without quantization
Free compiler
Compile models to .tic
We email your activation key as soon as you submit. Compile locally, your models stay in your environment.