AI: LLM Spend Control with LiteLLM Budgets and OpenAI Limits
0 (0)

24 September 2026

The main goal we had when introducing LiteLLM AI Gateway on the project (see LiteLLM: AI Gateway for LLMs – features overview) was access and cost control, because right now we can’t enable auto-recharge for our OpenAI or OpenRouter account: our startup is still in the development and experimentation stage, developers actively use AI agents… Read More: AI: LLM Spend Control with LiteLLM Budgets and OpenAI Limits0… »

Loading

LiteLLM: Debugging AI Cost Monitoring with VictoriaMetrics
5 (1)

16 September 2026

Had a pretty interesting case with monitoring LLM costs through LiteLLM when using multiple providers. What we have: LiteLLM: AI Gateway, all our services work through it, it proxies requests to providers, generates metrics and traces, controls access, spending, etc. OpenAI and OpenRouter: currently the two main providers LiteLLM sends requests to some clients (our… Read More: LiteLLM: Debugging AI Cost Monitoring with VictoriaMetrics5 (1) »

Loading

LiteLLM: Custom Callbacks and LLM Evaluations with Judge LLM
0 (0)

11 September 2026

A quick recap of what we are doing on the project right now: we have a separate hardware server (hostname == “Matrix”), where we run our own self-hosted models with llama.cpp. In our Kubernetes cluster we have a LiteLLM AI Gateway for our clients – Backend API and other project services. Clients send their OpenAI/Anthropic/OpenRouter… Read More: LiteLLM: Custom Callbacks and LLM Evaluations with Judge LLM0 (0) »

Loading

LiteLLM: Custom Callback for Traffic Mirroring and OTel Tracing to VictoriaTraces
0 (0)

9 September 2026

We’re currently setting up our own hardware server where we want to run self-hosted models. But we can’t just switch client traffic to them right away – first we need to see how these self-hosted LLMs will actually perform. So the general idea for now is to keep sending traffic to the primary provider, OpenAI… Read More: LiteLLM: Custom Callback for Traffic Mirroring and OTel Tracing to… »

Loading

LiteLLM: Traffic Mirroring, Batch Completions, and Traffic to Two Providers
0 (0)

21 August 2026

We’re getting ready to launch a self-hosted LLM, and at the testing stage the general idea is to send client requests simultaneously both to the “default production model” like GPT-5.6 and to the model running on our own server. And after getting the responses, we’ll compare them with Phoenix or Opik, and gradually tune our… Read More: LiteLLM: Traffic Mirroring, Batch Completions, and Traffic to Two Providers0… »

Loading

llama.cpp: Metrics and Monitoring with VictoriaMetrics
0 (0)

21 August 2026

We have a server where we’re going to run self-hosted LLMs. We spent quite a while choosing what exactly to use for running the models – vLLM, SGLang, or llama.cpp, and eventually settled on llama.cpp – at least for now. In the post NixOS: getting started, package installation, and system configuration I described installing Node… Read More: llama.cpp: Metrics and Monitoring with VictoriaMetrics0 (0) »

Loading

Ubuntu: Installing NGINX with TLS Certificate from Let’s Encrypt and AWS Route 53
0 (0)

17 August 2026

We have an Ubuntu server running one of our services. The service itself is still at the Proof of Concept stage, so we’re not moving it to Kubernetes yet – but sending traffic in plain text, especially when these are requests to AI Agents with prompts and responses, is not a great idea. So, the… Read More: Ubuntu: Installing NGINX with TLS Certificate from Let’s Encrypt and… »

Loading

LiteLLM: Metrics, Traces, and Debugging exception_class=”ValueError”
0 (0)

12 August 2026

A few days ago, I ran into an interesting situation with LiteLLM: on the one hand, the metrics showed a lot of errors “from the provider”, while on the other hand, the traces and alerts showed only a single error. I had to dig into it a bit and figure out some nuances of how… Read More: LiteLLM: Metrics, Traces, and Debugging exception_class=”ValueError”0 (0) »

Loading

NixOS: Getting Started, Installing Packages, and Configuring the System
0 (0)

5 August 2026

We got a new instance, a hardware server that will run our self-hosted LLMs. The server will run NixOS – not my choice, but the system looks interesting. I’ve been hearing about it for a long time, and now I have a great opportunity to get familiar with it. For now, my part is only… Read More: NixOS: Getting Started, Installing Packages, and Configuring the System0 (0) »

Loading

LiteLLM: OpenRouter Integration and Fallbacks Configuration
0 (0)

31 July 2026

OpenRouter recently announced a 50% discount on OpenAI, so we decided to give it a try. We’ll switch things over on LiteLLM, which already handles requests from all our services – so the switch should be pretty simple. Before OpenRouter, we’ll set up fallbacks straight to OpenAI in LiteLLM – for cases when requests to… Read More: LiteLLM: OpenRouter Integration and Fallbacks Configuration0 (0) »

Loading