vLLM is the high-throughput, GPU-accelerated open-weight server most teams reach for once a single laptop stops being enough. Continuous batching, prefix caching, multi-GPU sharding, OpenAI-compatible HTTP. Run it on your own boxes or through providers like Together. Pair with Digitorn the same way you pair with Ollama, the YAML does not care.
In Studio, choose vLLM as your agent's brain from the model picker, or just ask the builder assistant for it. There's no config to write, it's wired up for you.
vLLM doesn't have a first-class connector yet. You can still use it as your agent's brain, and native one-click support is on the roadmap.
VLLM_API_KEYOptional auth token if you front the server with an auth layerOpen Studio, connect vLLM, and tell the builder assistant what you want. It builds the agent for you and you refine it live on the canvas. It's vibe coding for agents, no setup to run.
Open StudioEngineering notes from the Digitorn team. No marketing, no launch announcements, no "10 prompts that will change your life". Just the things we write that we'd want to read.