Thanks to visit codestin.com
Credit goes to github.com

Skip to content

Expose per-model memory usage (the per-process VRAM walk already computes it, then discards the split) #11970

Description

@YourYoungerBrothersPug

Is your feature request related to a problem? Please describe.
There is no way to ask LocalAI how much memory each loaded model holds, so a
downstream scheduler cannot split a shared GPU into "what the engine holds" and
"what everything else holds".

I run chat, embeddings, transcription and image generation through one LocalAI
container on one consumer card, and admit background work only when there is room
left for a user's request. Without a per-model figure that subtraction is
unavailable and background work has to be admitted unweighed.

Checked on v4.9.0 (f7ad3f7), image
localai/localai:v4.9.0-gpu-nvidia-cuda-12, RTX 3060, two models resident:

  • GET /system — ids and backend names only, no figures.
  • GET /api/pssize and size_vram are literal 0. Filed separately as a bug.
  • GET /backend/monitor?model=granite-4.1-8b — 500 here, and gopsutil RSS/VMS
    is host process memory rather than device memory even when it works:
    rpc error: code = Unimplemented desc = , then failed local node process sample: no grpc backend found for llama-cpp/models/granite-4.1-8b/granite-4.1-8b-Q4_K_M.gguf.bin
  • GET /api/resources — device aggregate and per-GPU totals, which is genuinely
    useful and I'll switch to it, but it cannot attribute usage to a model.
  • POST /api/models/vram-estimate — a prediction from weight files, not a reading.
    It can't see the actual offload split, and being wrong in the optimistic
    direction is what crashes the box.
  • GET /metrics — Go runtime metrics only.

Describe the solution you'd like
Per-model resident memory on /system, as optional per-entry fields:

{"loaded_models":[{"id":"granite-4.1-8b","backend":"llama-cpp","size_vram":5100000000}]}

Omitting the field (or null) where a backend can't report it is much better than 0,
for the reason in the linked bug. Weights alone would already be a large improvement
over nothing; a KV-cache split would be ideal but isn't required. /system vs. a new
endpoint vs. populating /api/ps — no preference, whichever fits the API surface.

Most of the machinery looks like it's already there on Linux. drmFdInfoUsageByRenderNode
walks /proc/[0-9]*/fd for DRM render fds and reads drm-total-<region> per PID:

https://github.com/mudler/LocalAI/blob/v4.9.0/pkg/xsysinfo/drmfdinfo.go#L48-L78

but then does out[renderName] += ..., aggregating per GPU and dropping the per-PID
breakdown. Since each backend is its own gRPC process, keeping the PID key and joining
it against the loaded-model list would give per-model VRAM with no new probing. That
path covers amdgpu/i915/xe on kernel >= 5.19 and not NVIDIA's proprietary driver, which
doesn't emit those keys — which is exactly why I'm asking for an optional per-entry
field rather than a guaranteed one. Please don't let uneven backend support block it:
partial coverage that says so is far more useful than none.

Describe alternatives you've considered

  • nvidia-smi --query-compute-apps — can't work across vendors, which is a large part
    of why LocalAI is attractive here. It also returns no rows at all in-container on a
    WSL2 host, while the card correctly reports 10849/12288 MiB used.
  • Summing /backend/monitor — rejected even with the 500 fixed: host RSS/VMS says
    nothing about VRAM and double-counts mapped weights.
  • Trusting /api/ps — would over-commit the card, per the linked bug.
  • Aggregate device reading plus a measured "everything else" peak — what I do today.
    It works but is coarse: the engine's own share is never known.
  • Estimating from the GGUF plus a modelled KV cache, i.e. what vram-estimate does —
    a guess wearing a measurement's clothes, and wrong in the direction that crashes.

Additional context
Searched existing issues for size_vram, api/ps, GPUMemoryInfo and vram-in-title and found nothing covering this. I'm not able to take the PR, but happy to test a change against this setup and to report what each backend does or doesn't report.

Written by my beloved Claude Code :)

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions