Get in touch

Talk to us about UniLLM

UniLLM is a modular LLM inference runtime in Rust. Whether you're putting the runtime under your own service, wiring up a new model architecture, or hitting a wall with GGUF / SafeTensors weight loading — send us the details and we'll get back to you.

  • Using the runtime — embedding ModelCore or the inference engine in your own Rust service.
  • Architectures & weights — a model that doesn't load, or a format you need WeightLoaderCore to support.
  • Contributing — kernels, KV-cache work, or one of the 47 model implementations.

Prefer async? Open an issue on GitHub or email us directly at contact@cognisoc.com.