Tagged with

Ai

This post thumbnail

18 July 2026 07:30 PM

vLLM, SGLang, TGI, Triton, and TensorRT-LLM solve different problems. Here is how to pick the right serving engine for an LLM workload on Kubernetes, with a real engine-swap walkthrough.

This post thumbnail

24 June 2026 06:30 PM

Deploy Qwen2.5-1.5B-Instruct on a Kubernetes GPU node with vLLM, expose it as an OpenAI-compatible API, and verify it with a real curl request.

This post thumbnail

7 May 2026 07:49 AM

The best use of AI in DevOps isn't autonomous agents with production access. It's reducing cognitive toil: reading docs, summarizing release notes, comparing configs, and giving engineers enough context to make better decisions.