Basically anything llama.cpp (Vulkan backend) should work out of the box w/o muc...

lhl 53 days ago | parent | context | favorite | on: Show HN: Real-time AI Voice Chat at ~500ms Latency

Basically anything llama.cpp (Vulkan backend) should work out of the box w/o much fuss (LM Studio, Ollama, etc).

The HIP backend can have a big prefill speed boost on some architectures (high-end RDNA3 for example). For everything else, I keep notes here: https://llm-tracker.info/howto/AMD-GPUs