Skip to content

Product

ISIRO Runtime

Smaller weights · More KV cache

Install

curl -fsSL https://isiro.ai/install.sh | sh

Paste in terminal

Production deployment? Start a pilot

Overview

Efficiency without loss

Compile models once into .tic with smaller memory footprint, then deploy through ISIRO Runtime, which integrates the inference frameworks you already use.

Supports BF16 vLLM on NVIDIA GPUs today. More precisions, frameworks, and hardware are in progress.

Technical Overview

Measured results

  • Up to 30%

    Smaller weights, more KV cache

  • Exact

    Bit-exact weights without quantization

Sample models & benchmarks

Free compiler

Compile models to .tic

We email your activation key as soon as you submit. Compile locally, your models stay in your environment.

ISIRO EULA *