inference-server
Here are 107 public repositories matching this topic...
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the macOS menu bar
-
Updated
Jun 18, 2026 - Python
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
-
Updated
Jun 17, 2026 - Rust
RamaLama is an open-source developer tool that simplifies the local serving of AI models from any source and facilitates their use for inference in production, all through the familiar language of containers.
-
Updated
Jun 17, 2026 - Python
Turn any computer or edge device into a command center for your computer vision projects.
-
Updated
Jun 18, 2026 - Python
Open-source inference server and production cluster for all the models your agent needs.
-
Updated
Jun 16, 2026 - Python
The simplest way to serve AI/ML models in production
-
Updated
Jun 18, 2026 - Python
An open-source computer vision framework to build and deploy apps in minutes
-
Updated
May 8, 2024 - Rust
Python + Inference - Model Deployment library in Python. Simplest model inference server ever.
-
Updated
Feb 14, 2023 - Python
A REST API for Caffe using Docker and Go
-
Updated
Jul 20, 2018 - C++
Fast GPU OCR server. 270 img/s on FUNSD. TensorRT FP16, PP-OCRv5, HTTP + gRPC.
-
Updated
Jun 11, 2026 - C++
PMetal: high-performance Apple Silicon framework for local LLM inference, LoRA/QLoRA fine-tuning, serving, quantization, and MLX/Metal acceleration.
-
Updated
Jun 5, 2026 - Rust
Work with LLMs on a local environment using containers
-
Updated
Jun 17, 2026 - TypeScript
This is a repository for an nocode object detection inference API using the Yolov3 and Yolov4 Darknet framework.
-
Updated
Jun 28, 2022 - Python
Auto-tuned launcher for GGUF models on llama.cpp / ik_llama.cpp — OpenAI-compatible server with multi-GPU tensor-split, MoE expert placement, measured flag tuning (AI Tune), hardware-matched HuggingFace downloads, and crash recovery. An Ollama alternative for multi-GPU rigs.
-
Updated
Jun 18, 2026 - Go
This is a repository for an nocode object detection inference API using the Yolov4 and Yolov3 Opencv.
-
Updated
Jun 28, 2022 - Python
ONNX Runtime Server: The ONNX Runtime Server is a server that provides TCP and HTTP/HTTPS REST APIs for ONNX inference.
-
Updated
May 10, 2026 - C++
This is a repository for an object detection inference API using the Tensorflow framework.
-
Updated
Jun 28, 2022 - Python
Serving AI/ML models in the open standard formats PMML and ONNX with both HTTP (REST API) and gRPC endpoints
-
Updated
Feb 24, 2026 - Scala
Orkhon: ML Inference Framework and Server Runtime
-
Updated
Feb 1, 2021 - Rust
Deploy DL/ ML inference pipelines with minimal extra code.
-
Updated
Feb 10, 2026 - Python
Improve this page
Add a description, image, and links to the inference-server topic page so that developers can more easily learn about it.
Add this topic to your repo
To associate your repository with the inference-server topic, visit your repo's landing page and select "manage topics."