Janus runs GGUF models via Vulkan on AMD, Intel, and Nvidia GPUs without CUDA
Open-source project in Go published as a Show HN uses llama.cpp's Vulkan backend to run GGUF models locally on any GPU, with an OpenAI-compatible API and no need for Python, Docker, or Ollama.
What Janus is
Janus is an open-source project published on Hacker News as a Show HN by the developer behind the Vibra-Ingenn/Janus repository on GitHub. The idea is straightforward: a single binary written in Go that loads models in the .gguf format (the format used by llama.cpp) and exposes an API compatible with OpenAI's, running entirely on the user's machine.

No Python, Docker, or Ollama installation is required. The repository has an MIT license, two commits of history, and 13 stars as of this writing: a project fresh out of the oven, not a mature tool with years of adoption.
Among the features listed in the README are model switching without restarting the server (/models/load), support for the "thinking" model pattern (which separates reasoning in <think> tags from the response content), and automatic detection of the chat template from the GGUF file's own metadata.
The differentiator: Vulkan instead of CUDA
Most local inference tools that run on GPU depend on CUDA, which in practice restricts hardware acceleration to Nvidia cards. That's the case for much of the local LLM server ecosystem, which uses direct bindings to CUDA or vendor-specific ROCm builds.
Janus uses llama.cpp's Vulkan backend. Vulkan is a cross-platform graphics and compute API maintained by the Khronos Group, with drivers for AMD, Intel, and Nvidia GPUs (including integrated ones). In practice, this means the same binary can accelerate inference on laptops with integrated Intel GPUs, desktops with AMD cards, or machines with Nvidia, without switching runtimes.
When there's no compatible GPU, Janus falls back to CPU. On macOS, Vulkan support "varies by hardware," according to the README itself, and the recommended path there is the CPU backend.
How to install and get it running
On Windows, which the project treats as its primary platform, the flow is to clone the repository and run the build script, which downloads precompiled llama.cpp DLLs with Vulkan support:
git clone https://github.com/Vibra-Ingenn/Janus.git
cd Janus
.\build.ps1Then you need to place a .gguf file in the models/ folder, either by downloading it manually from Hugging Face or using the downloader included in the repository itself, modelget:
go build -o dist\modelget.exe .\cmd\modelget
.\dist\modelget.exe -repo meta-llama/Llama-3.2-3B-Instruct -file Llama-3.2-3B-Instruct-Q8_0.gguf -out .\models\Configuration lives in a .env file, copied from .env.example. The main variables are INFERENCE_BACKEND (vulkan, cpu, or openrouter), JANUS_MODEL_PATH (the .gguf path), and JANUS_GPU_LAYERS, which defaults to -1 to push all of the model's layers onto the GPU. On Linux and macOS the process is the standard go build, with the caveat that Linux needs libllama.so in the same directory as the binary or in LD_LIBRARY_PATH.
An API that already speaks your tools' language
The server starts by default on 127.0.0.1:8990 (not on port 8080, as the README explicitly warns) and exposes the following paths:
| Method | Route | What it's for |
|---|---|---|
| GET | /health | process liveness check |
| GET | /v1/models | list of models in the OpenAI format |
| POST | /v1/chat/completions | chat, with streaming support |
| POST | /models/load | model switching without restarting |
| GET | /models/list | .gguf files available locally |
| GET | /engine/status | VRAM usage and active backend |
Since the API follows the OpenAI contract, the same endpoint serves curl, scripts, and tools like Cursor or Cline, as long as you point the base URL to http://127.0.0.1:8990/v1 and leave the API key field blank. This removes the step of rewriting integrations that already exist for the OpenAI API just to test a local model.
Hot-swap and "thinking" models
Switching models without taking down the server is a concrete need for anyone testing several .gguf files in the same day: now it means calling /models/load with a different file path, instead of killing the process, editing the .env, and starting it back up.
Support for explicit-reasoning models, which emit the thinking step between <think> tags, is automatically split by Janus into a reasoning_content field in the response, keeping the final content isolated from the intermediate reasoning. It's the same kind of separation that models like the DeepSeek-R1 family popularized, and that chat tools need to handle differently from the response text.
Janus vs. Ollama vs. the llama.cpp server
The niche of "running GGUF locally with an HTTP API" already has established competition: Ollama and llama.cpp's own built-in server have covered the same use case for longer, with larger communities and model libraries ready to download with a single command.
Janus's point of difference lies in the simplicity of its distribution (a single .exe on Windows, with no external runtime) and in its specific bet on Vulkan as a universal acceleration path, instead of relying on separate builds per GPU vendor. That's a reading of positioning, not a data point from the project: the README doesn't include a benchmark comparing the three.
What remains open
The project doesn't publish performance numbers, nor a speed comparison between the Vulkan backend and native CUDA on the same Nvidia GPU. With 13 stars and two commits of history at the time of publication, there are also no tagged releases on GitHub yet, which makes production use premature.
For anyone who just wants to test local models without getting stuck on CUDA drivers, especially on a machine with an AMD or Intel GPU, Janus is a path simple enough to get running in an afternoon. To decide whether it replaces Ollama or the llama.cpp server in your workflow, the project still needs to mature and the community needs to test it across more hardware combinations.
Translated from the Brazilian Portuguese original · Read the original
SoftBank completes $30 billion investment in OpenAI, reaching 13% stake in the company
The Japanese giant closed the final installment of a $30 billion commitment, raising its total investment in ChatGPT's creator to $64.6 billion and its stake in the company to 13%.