LLM28 Sept 2026· 1 min readEvaluating LLM output without vibesBuild a small eval harness before you touch the prompt.#llm#evals#testing
RAG24 Sept 2026· 1 min readRAG basics: chunking and retrievalWhy most RAG failures are retrieval failures.#rag#embeddings
Performance18 Sept 2026· 1 min readPrompt caching and latencyCut cost and time-to-first-token by reusing the prefix.#llm#caching#latency