A powerful, Windows-native GUI launcher for llama.cpp's llama-server.exe. Streamline model management, server configuration, and local LLM inference with an intuitive interface, automatic GGUF scanning, and a robust caching system.
Platform: Windows 10/11 (64-bit) | Language: Delphi/Pascal
| Category | Description |
|---|---|
| 📦 GGUF Management | Automatic multi-threaded directory scanning, shard validation, and metadata extraction (architecture, layers, context, size, embedding, etc.) |
| 💾 Smart Cache | Persistent models-cache.txt for instant metadata retrieval. Avoids repeated file parsing on startup |
| ⚙️ Server Configuration | Full control over llama-server parameters: threads, context, GPU/CPU offloading, sampling (temp, top_k/p, min_p, repeat/presence penalty), flash-attn, KV cache types, and advanced flags |
| 📋 Preset System | Per-model presets (auto-saved on launch) & generation presets (temperature, reasoning mode, seed, etc.) stored as .ini files |
| 🌍 Multi-Language | 25+ supported languages via INI-based translation system. Fallback to EN/FR if translation missing |
| 📊 Real-time Logging | High-performance log buffer, memo display, auto-scroll, file export, and max-lines limiter |
| 🔍 System Monitor | Live CPU/RAM/VRAM tracking in the status bar |
| 🖥️ UX Enhancements | Tray icon with minimize-to-tray, drag & drop GGUF support, single-instance mutex, WebUI launcher, known arguments helper |
llama-server.exe (from llama.cpp)💡 The launcher is self-contained. All Delphi RTL and Indy components are statically linked or bundled.
LlamaLaunchers.zip) from the Releases page.C:\Tools\LlamaLauncher\).llama-server.exe in a known directory (e.g., C:\llama.cpp\).LlamaLaunchers.exe.🔧 Settings → llama.cpp Path..gguf model (via Browse, History, or Drag & Drop).LlamaLauncher/
├── LlamaLaunchers.exe
├── presets/
│ ├── models-cache.txt # Cached GGUF metadata
│ ├── models/ # Per-model presets (.ini)
│ └── gen/ # Generation presets (.ini)
├── lang/ # Language translation files (.ini)
└── logs/ # Runtime logs (.log)
.gguf file onto the window.-00001-of-00003.gguf pattern), aggregates sizes, and extracts GGUF metadata.fp16, q8_0, etc.), Rope Scaling, Batch/Ubatch sizes, Tools/Props flags, Extra CLI arguments.File → Load Preset.🧠 Gen Presets combobox. Save/Overwrite/Delete directly from the UI.Restore.Translations are stored in lang/. The app auto-detects available languages and falls back to EN or FR if a translation is missing.
To add a new language:
lang/.ini in the app directory.[FormName.ComponentName] sections with Caption and Hint keys).| Issue | Solution |
|---|---|
llama-server.exe fails to launch or shows DLL errors |
Install the Visual C++ 2015-2022 Redistributable (x64). Specifically, VC14_redist.x64-14.50.35719.exe (or a compatible version) is required for llama-server.exe to run correctly. Download it from Microsoft's official site if missing. |
GGUF not detected or missing shards |
Ensure all shard files (-00001-of-NNNNN.gguf) are present in the same directory. The scanner validates completeness. |
| Server fails to start | Verify llama-server.exe path, check logs/llama_launcher_log_*.log for CLI errors, and ensure required dependencies (CUDA/ROCm) are installed. |
| Cache outdated | Delete presets\models-cache.txt or click 🔄 Update in the model info panel. |
| UI language missing | Ensure lang/ |
| High memory usage during scan | The scanner is multi-threaded but reads sequentially. Use Fast Scan mode in settings to skip tensor info parsing. |
This project is distributed under the MIT License. See LICENSE for details. llama.cpp and its ecosystem are developed by Georgi Gerganov and the open-source community under the Apache 2.0 License.
Built with ❤️ for the local AI community. Run powerful models locally, effortlessly.