Content-Length: 585975 | pFad | http://github.com/topics/multi-gpu

ad multi-gpu · GitHub Topics · GitHub
Skip to content
#

multi-gpu

Here are 173 public repositories matching this topic...

Kubernetes operator for self-hosted LLM inference across a heterogeneous GPU fleet: NVIDIA CUDA, AMD Vulkan, and Apple Silicon Metal. Runtimes: llama.cpp, vLLM, TGI, mlx-server. Multi-GPU sharding, model caching, OpenAI-compatible endpoints. Apache-2.0, run across homelab and on-prem fleets, actively developed.

  • Updated Aug 3, 2026
  • Go

Improve this page

Add a description, image, and links to the multi-gpu topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the multi-gpu topic, visit your repo's landing page and select "manage topics."

Learn more









ApplySandwichStrip

pFad - (p)hone/(F)rame/(a)nonymizer/(d)eclutterfier!      Saves Data!


--- a PPN by Garber Painting Akron. With Image Size Reduction included!

Fetched URL: http://github.com/topics/multi-gpu

Alternative Proxies:

Alternative Proxy

pFad Proxy

pFad v3 Proxy

pFad v4 Proxy