Thanks to visit codestin.com
Credit goes to github.com

Skip to content
#

inference-embedded-engine-models

Here is 1 public repository matching this topic...

Hardware probe and best-of-N sampling harness in C. Ranks GGUF quantisations against measured VRAM and RAM, drives Ollama or llama.cpp over OpenAI-format HTTP on localhost, and gates every candidate behind a static check and a test run.

  • Updated Sep 6, 2026
  • C

Add this topic to your repo

To associate your repository with the inference-embedded-engine-models topic, visit your repo's landing page and select "manage topics."

Learn more