Hardware probe and best-of-N sampling harness in C. Ranks GGUF quantisations against measured VRAM and RAM, drives Ollama or llama.cpp over OpenAI-format HTTP on localhost, and gates every candidate behind a static check and a test run.
ai inference inference-server ai-agents inference-engine inference-optimization inference-acceleration inference-api ai-model ai-tools ai-agent inference-embedded-engine inference-embedded-engine-models ai-coding inference-gateway inference-providers
-
Updated
Sep 6, 2026 - C