Quick answer: a 16GB card runs a 7B to 14B model at Q6_K with a comfortable context window. The build takes an evening, and the payoff is a model that answers offline and never phones home.
Running a language model on your own hardware keeps growing in appeal, and for good reason. A local setup answers questions after hours, never phones home, and keeps client data on the bench instead of in somebody else's cloud. Here is what actually works on an RX 9070 XT with 16GB of memory, without the marketing.
What 16GB of VRAM actually buys you
The card's memory sets the boundary for everything else. A quantised model must fit alongside the context window you plan to use. Most people on this class of card run 7B to 14B parameter models, and the sweet spot for quality per gigabyte is a Q6_K quant. Roughly:
- A 7B model at Q6_K sits around 6GB, leaving plenty of room for a long conversation.
- A 14B model at Q6_K comes in near 11GB, which fits but keeps the context window shorter.
- Bigger models are possible at lighter quants, but quality drops and the trade usually is not worth it.
Building llama.cpp
Start with llama.cpp. Clone the repository, install the build prerequisites, and compile with the Vulkan backend enabled. On RDNA4 cards Vulkan runs a model comfortably out of the box, and it keeps the build simple and portable. ROCm works too, but Vulkan is the path of least resistance here. The build produces two tools, llama-cli and llama-server, which is all you need to test.
Picking a model and talking to it
Grab a Q6_K quant from a repository you trust and point llama.cpp at it. Q6_K runs close to the quality of the full precision model while fitting comfortably in memory, and on this card speed stays in the range where conversation feels natural rather than like waiting for a download.
Keep an eye on memory during long sessions. Context length is the second consumer of VRAM, and long conversations push memory use upward. If a session starts to crawl, the usual cause is the context window growing past the budget, not the model itself.
Making the model useful to other programs
Run server mode instead of the command line when you want other applications on the same machine to reach the model. It exposes a simple API on your own network, and tools can point at it the same way they point at any hosted service. The difference is the server never leaves the building.
The honest summary
A 16GB card runs a small model comfortably, and the experience of a model on your own hardware is different in kind from a rented endpoint. It will not run the largest models, and it will not beat a data centre. What it gives you is a private assistant that works after hours, offline, for the price of electricity.
Not sure whether your machine is up to it? Bring it in and we will check the hardware and set expectations honestly, before you spend a cent on parts.