FAQ
Frequently Asked Questions
No. Weights stay bit-exact. Memory savings come from lossless compact .tic and fused serve. Weight precision stays the same.
No. Weights are not quantized, approximated, retrained, or calibrated down. Model output remains the same.
Decoded .tic weights match the raw model weights bit-for-bit, with no loss or precision change. A hash manifest verifies that equality.
Moving less data on the GPU lowers energy and infrastructure cost. The same GPUs can do more work, or you can right-size instances for the same model.
See Models & Benchmarks for representative footprint results, and the linked GitHub benchmarks for memory, latency, and throughput.
Install for evaluation, and use the free compiler to build your own .tic. For production or enterprise deployment, start a pilot.
Self-managed and self-hosted cloud, on-prem, and local AI setups where you run inference. BF16 with vLLM on NVIDIA GPUs.