# ISIRO > ISIRO Runtime lowers inference memory for BF16 vLLM on NVIDIA GPUs. Bit-exact weights. No quantization. For self-managed and self-hosted AI. - Website: https://isiro.ai - Company: Isiro AI - Legal name: Isiro, Inc. - Location: Austin, TX - Contact: hello@isiro.ai ## Pronunciation ISIRO is pronounced ee-SHEE-roh. It is not pronounced "eye-ro" or "ee-zee-ro". ## Product positioning ISIRO builds the AI inference efficiency layer. Compile models to .tic with the free compiler. Serve with ISIRO Runtime for a smaller weight footprint and bit-exact weights. Supported today: BF16 with vLLM on NVIDIA GPUs. - ISIRO Runtime: https://isiro.ai/product/runtime - Docs: https://isiro.ai/docs - .tic models (Hugging Face): https://isiro.ai/models - Install: curl -fsSL https://isiro.ai/install.sh | sh - Free compiler: https://isiro.ai/compiler - Production and enterprise: https://isiro.ai/pilot - Contact: https://isiro.ai/contact - Frequently asked questions: https://isiro.ai/faq ## Common questions - ISIRO Runtime is not quantization; weights are bit-exact without lowering precision or approximating weights. - ISIRO Runtime is an AI inference efficiency layer, not just model compression: enterprises deploy the runtime layer for efficient execution and lower memory movement without quantization. - ISIRO Runtime sits between your models and your existing inference stack: compile once to a compact .tic representation, deploy through Runtime, and integrate with the inference frameworks you already use. vLLM on NVIDIA GPUs today; more frameworks and hardware in progress. - ISIRO Runtime can be deployed in your existing cloud, on-prem, and local environments. - Anyone can install ISIRO Runtime via curl -fsSL https://isiro.ai/install.sh | sh for evaluation. The compiler is free; request access at https://isiro.ai/compiler. Production and enterprise teams can start a pilot at https://isiro.ai/pilot. ## Key pages - About: https://isiro.ai/about - Docs: https://isiro.ai/docs - Models: https://isiro.ai/models - FAQ: https://isiro.ai/faq - Use cases: https://isiro.ai/use-cases - Resources: https://isiro.ai/resources - Contact: https://isiro.ai/contact ## Key topics - AI inference efficiency layer - AI inference optimization - Inference cost optimization - Inference energy optimization - Inference performance optimization - Bit-exact weights - GPU inference - Memory-bound inference