One machine, many models
Running one model on your own computer is the easy part. Running several, for several people, is where it comes apart: two clients collide on the same server, a picture job takes the memory the language model was holding, every engine wants its own environment, and none of it was built to be shared.
llamanager is the layer that makes one computer behave like a service. It installs the engines, downloads the weights, supervises the processes, and puts a single queue in front of all of them. Text, image, video, and speech reach the same address, pass the same door, and show up on the same dashboard.
We built it for our own work and we run it every day. It is private software, offered on a quote: we install it on your hardware, choose the models that suit the work, and stay on hand afterwards. What you generate stays on the machine that generated it.
- Four familiestext, image, video, speech
- One addressin the formats your tools already speak
- Your machinemacOS, Linux, Windows
Four families, one queue
Every family installs from the same page and is served through the same queue.
Text generation
Language models behind an address your existing tools already speak. Name a different model in a request and the queue swaps it in without dropping work in flight, or keep several warm side by side, each in its own slot, so every request finds the one it asked for with no waiting.
Image generation
Several image engines, installed from the dashboard instead of assembled by hand. Generate from a prompt, or from reference pictures for editing, composition, and variations. Every result lands in a gallery on disk beside the settings that made it.
Video generation
Clips from a prompt, from an opening frame, or from a handful of reference stills. Before anything loads, llamanager works out what the clip will cost in graphics memory and says plainly that it will not fit, naming a size that will, instead of dying halfway through.
Audio generation
Sound has its place in the same queue as everything else. Generated clips can carry a soundtrack made alongside the picture, and audio work is installed, queued, and served the same way as every other family, from the same dashboard.
Speech transcription
Spoken audio to text, with timing and confidence for every word, from a file or live from a microphone. Transcription keeps a worker warm and serves many requests at once inside a memory budget, so it runs beside a loaded language model instead of evicting it.
Built to be shared
A queue with priorities
Every request joins one queue, tagged with the client it came from and the priority that client was given. A long overnight batch cannot block the editor you are typing in, and cancelling a job stops the work itself, whether it is waiting or already running.
People and keys
Each person and each tool gets its own key, kept only as a hash. Priorities are set per key, and any key can be switched off and back on, so a client can be parked for a week without rotating or deleting anything. The activity feed records who asked for what.
Engines and weights, installed
Engines install from a button. llamanager builds the environment each one needs, keeps them apart so they cannot break each other, and shows the disk it costs and the log as it goes. Open weight models are pulled from public repositories the same way, whole or in part, with progress on the page.
One graphics card, shared
Families that would otherwise fight over the same memory take turns by default and share it when there is headroom, on a policy you set. Transcription runs inside a budget that leaves the language model where it is. A crashed engine is restarted, with a cap so a broken model cannot spin forever.
One address for everything
All of it answers on a single endpoint, in the formats the common client libraries already use, so existing code points at your machine by changing one line. Assistants and agents can drive the machine as well as talk to it: list models, watch memory, pull weights, generate, transcribe.
Glance at it from your phone
The dashboard is a real phone surface, not a shrunken desktop. Over your own private network you can see the queue moving, the model loaded, and what crashed while you were out, without opening a terminal.
Where it runs
- Operating systems
- macOS, Linux, and Windows. It can start with the machine or at login, and it is managed the same way on all three.
- Graphics
- Cards from Apple, NVIDIA, AMD, and Intel are detected and their memory reported live. Without a supported card it still runs on the processor, more slowly.
- Reaching it
- It listens only to the machine itself until you say otherwise. Open it to your own private network and the dashboard and the endpoint follow you to a phone or another laptop.
- What leaves
- Prompts, pictures, clips, and recordings stay where they were made. The only traffic llamanager starts is fetching the engines and weights you asked for, and checking whether a newer version exists.
Common questions
What is llamanager?
llamanager is our software for running models on a machine you own. It installs the engines, downloads the model weights, and serves text, image, video, and speech work through one queue, one address, and one dashboard, with a key and a priority for each person using it.
Does our data leave the machine?
What you generate stays on the machine that generated it. Prompts, pictures, clips, and recordings are written to its own disk and are not sent anywhere. The only outbound traffic llamanager starts is downloading the engines and model weights you ask for, and checking whether a newer version of itself exists.
What hardware does it need?
A computer with a graphics card is the comfortable case, and the more memory that card has, the larger the models and the longer the clips it can carry. There is no hard minimum: llamanager reads what the machine has, suggests model sizes that fit, and refuses work that would not, naming a size that would. Tell us the machine you have and we will tell you what it can run.
Can several people share one machine?
Yes, and that is the reason it exists. Each person and each tool has its own key and its own priority, everything lands in the same queue, and one client's long job cannot starve another. Any key can be switched off without deleting it.
Does it work with the tools we already use?
Yes. It answers in the formats the common client libraries already speak, so code written against a hosted service usually needs nothing but a new address and key. Coding assistants and agents can also drive the machine itself: loading models, watching memory, generating, transcribing.
How do we get llamanager?
llamanager is private software and is not sold as a download. We install it on your hardware, set it up for the people and tools that will use it, choose the models that suit the work, and stay on hand afterwards. Each setup is priced on a quote: tell us what you want to run and on which machine.
Run it on your own machine
Tell us what you want to generate, transcribe, or keep in house, and what hardware you have. We will tell you what llamanager can do with it, and what it costs.
- What you want to run: chat, images, video, transcription
- The hardware you have, or plan to buy
- How many people and tools will share it
Opens your email with these points ready to fill in.