Host-aware recommendations
Read RAM/VRAM/thermals and model size/quant hints; propose n_gpu_layers, context, and batch envelopes.
Under Validation
A local MCP advisor for developers already running GGUF models. Profiles hardware and suggests (or applies) offload, context, and batch settings for llama.cpp and Ollama. Not a new inference engine.
Concept validation only. No payment, no commitment, no recurring newsletter.
Proposed workflow
Advisor first. Hot-reload of every backend parameter is not assumed; some runners still need a restart.
Read RAM/VRAM/thermals and model size/quant hints; propose n_gpu_layers, context, and batch envelopes.
Expose tune/recommend tools to agents already in the developer workflow instead of a separate dashboard.
No telemetry of prompts or model weights. Runs beside your existing Ollama or llama.cpp install.
The honest status
The concept is being validated before development time is committed. Joining tells us the problem is relevant to you and gives you first access if the evidence supports a build.
Before you decide
The current scope, privacy model, and next step without launch-day promises.
No. It sits beside them. If your GUI already picks safe defaults, you may not need it.
Not guaranteed. Some parameters require a model reload. The product hypothesis includes honest limits in the UX.
No. The concept should state its limits clearly and preserve human review wherever judgment is required.
Your email is recorded for this experiment only. You receive one relevant update or beta invitation if the concept moves forward.
Early access
Join the waitlist. Planned pricing around a €19 one-time license or low monthly if demand clears the bar — no charge today.
Leave your email to receive the beta invitation if this concept moves forward.