🦙 Llama Launcher

A powerful, Windows-native GUI launcher for llama.cpp's llama-server.exe. Streamline model management, server configuration, and local LLM inference with an intuitive interface, automatic GGUF scanning, and a robust caching system.

Platform: Windows 10/11 (64-bit) | Language: Delphi/Pascal

✨ Features

Category Description
📦 GGUF Management Automatic multi-threaded directory scanning, shard validation, and metadata extraction (architecture, layers, context, size, embedding, etc.)
💾 Smart Cache Persistent models-cache.txt for instant metadata retrieval. Avoids repeated file parsing on startup
⚙️ Server Configuration Full control over llama-server parameters: threads, context, GPU/CPU offloading, sampling (temp, top_k/p, min_p, repeat/presence penalty), flash-attn, KV cache types, and advanced flags
📋 Preset System Per-model presets (auto-saved on launch) & generation presets (temperature, reasoning mode, seed, etc.) stored as .ini files
🌍 Multi-Language 25+ supported languages via INI-based translation system. Fallback to EN/FR if translation missing
📊 Real-time Logging High-performance log buffer, memo display, auto-scroll, file export, and max-lines limiter
🔍 System Monitor Live CPU/RAM/VRAM tracking in the status bar
🖥️ UX Enhancements Tray icon with minimize-to-tray, drag & drop GGUF support, single-instance mutex, WebUI launcher, known arguments helper

📦 Requirements

💡 The launcher is self-contained. All Delphi RTL and Indy components are statically linked or bundled.

📥 Installation & Setup

  1. Download the latest release (LlamaLaunchers.zip) from the Releases page.
  2. Extract the archive to your preferred directory (e.g., C:\Tools\LlamaLauncher\).
  3. Place your llama-server.exe in a known directory (e.g., C:\llama.cpp\).
  4. Launch LlamaLaunchers.exe.
  5. Configure the executable path via 🔧 Settings → llama.cpp Path.
  6. Select a .gguf model (via Browse, History, or Drag & Drop).
  7. Click ▶ Start to launch the server.

📁 Folder Structure (Auto-created)

LlamaLauncher/
├── LlamaLaunchers.exe
├── presets/
│ ├── models-cache.txt # Cached GGUF metadata
│ ├── models/ # Per-model presets (.ini)
│ └── gen/ # Generation presets (.ini)
├── lang/ # Language translation files (.ini)
└── logs/ # Runtime logs (.log)
    

🖥️ Usage Guide

🔹 Model Selection & Scanning

🔹 Server Configuration

🔹 Presets

🔹 Tray & Multi-Instance

🌍 Language Support

Translations are stored in lang/.ini. The app auto-detects available languages and falls back to EN or FR if a translation is missing.

To add a new language:

  1. Create lang/.ini in the app directory.
  2. Follow the existing INI structure ([FormName.ComponentName] sections with Caption and Hint keys).
  3. Restart the app.

🐛 Troubleshooting

Issue Solution
llama-server.exe fails to launch or shows DLL errors Install the Visual C++ 2015-2022 Redistributable (x64). Specifically, VC14_redist.x64-14.50.35719.exe (or a compatible version) is required for llama-server.exe to run correctly. Download it from Microsoft's official site if missing.
GGUF not detected or missing shards Ensure all shard files (-00001-of-NNNNN.gguf) are present in the same directory. The scanner validates completeness.
Server fails to start Verify llama-server.exe path, check logs/llama_launcher_log_*.log for CLI errors, and ensure required dependencies (CUDA/ROCm) are installed.
Cache outdated Delete presets\models-cache.txt or click 🔄 Update in the model info panel.
UI language missing Ensure lang/.ini exists. The app falls back to English/French automatically.
High memory usage during scan The scanner is multi-threaded but reads sequentially. Use Fast Scan mode in settings to skip tensor info parsing.

📜 License

This project is distributed under the MIT License. See LICENSE for details. llama.cpp and its ecosystem are developed by Georgi Gerganov and the open-source community under the Apache 2.0 License.

🤝 Credits & Acknowledgments

Built with ❤️ for the local AI community. Run powerful models locally, effortlessly.