AI SystemsAI Systems

How AI models are served and measured in practice: memory, batching, inference, caching, observability, and deployment.Cách mô hình AI được triển khai và đo lường trong thực tế: bộ nhớ, batching, inference, caching, observability và deployment.

14 articlesbài viết

  1. 01What One Model Call CostsWhy prefill and decoding have very different costs, and how to measure both phases on real hardware.
  2. 01Một lần gọi mô hình tốn gìVì sao prefill và decoding có chi phí rất khác nhau, cùng cách đo cả hai giai đoạn trên phần cứng thực tế.
  3. 02KV CacheEstimating KV-cache memory per token before running the model, then checking the formula against real measurements.
  4. 02KV cacheƯớc tính bộ nhớ KV cache cho mỗi token trước khi chạy mô hình, rồi kiểm tra công thức bằng số đo thực tế.
  5. 0312 GB: Out of Memory Before It Even RunsEstimating whether a model will fit on your GPU before downloading it, then validating the estimate against a real out-of-memory error.
  6. 0312 GB: chưa chạy đã hết chỗƯớc tính một mô hình có vừa GPU hay không trước khi tải xuống, rồi kiểm chứng phép tính bằng một lỗi hết bộ nhớ thực tế.
  7. 04When a Model Doesn't Fit on Your GPU: Three OptionsThree ways out when a model is bigger than the GPU (splitting across GPUs, offloading to RAM, quantizing), the cost of each on real hardware, and why the compute kernel matters as much as the bit width.
  8. 04Khi mô hình không vừa GPU: ba lựa chọnBa lối thoát khi mô hình lớn hơn card là chia qua nhiều GPU, đẩy sang RAM và lượng tử hóa, cái giá của từng lối trên phần cứng thật, và vì sao nhân tính toán quan trọng không kém số bit.
  9. 05Static BatchingWhy running many requests together raises throughput almost for free up to a predictable limit, what it costs in latency and memory, and why the simplest batching wastes nearly half its work.
  10. 05Gộp lô tĩnhVì sao chạy nhiều yêu cầu cùng lúc làm thông lượng tăng gần như miễn phí tới một giới hạn ước lượng được, cái giá về độ trễ và bộ nhớ, và vì sao gộp lô đơn giản lãng phí gần nửa công sức.
  11. 06Continuous Batching and Paged KV CacheSeparating the two ideas usually lumped together as "why vLLM is fast", designing a comparison that does not credit the whole gap to one idea, and why short requests end up waiting for long ones.
  12. 06Gộp lô liên tục và KV cache dạng trangTách hai ý tưởng hay bị gọi chung là "vì sao vLLM nhanh", thiết kế phép so sánh không quy hết chênh lệch cho một ý tưởng, và vì sao người hỏi câu ngắn phải chờ người hỏi câu dài.
  13. 07Measuring Serving Load CorrectlyWhy averages hide what users feel, how common benchmarking setups make latency look better than it is, and how to detect instability.
  14. 07Đo tải phục vụ cho đúngVì sao số trung bình che khuất trải nghiệm người dùng, cách benchmark phổ biến khiến latency trông tốt hơn thực tế, và cách phát hiện hệ thống mất ổn định.
  15. 08Inference Servers and StreamingWhat an inference server adds around the model, why it takes a minute to become ready, and measurements showing that streaming does not make the model any faster, it only changes when users start receiving output.
  16. 08Máy chủ suy luận và truyền theo luồngMáy chủ suy luận thêm gì quanh mô hình, vì sao mất cả phút mới sẵn sàng, và số đo cho thấy truyền theo luồng không làm mô hình nhanh hơn, chỉ đổi lúc người dùng bắt đầu nhận kết quả.
  17. 09Two Very Different Kinds of CacheA prefix cache reuses computation and does not change the answer, a semantic cache reuses old answers and can be wrong; measuring the first's benefit and why the second cannot be made safe with a cosine threshold alone.
  18. 09Hai loại bộ đệm rất khác nhauBộ đệm tiền tố tái dùng phép tính và không đổi câu trả lời, bộ đệm ngữ nghĩa tái dùng câu trả lời cũ và có thể trả sai; đo lợi ích của cái đầu và vì sao cái sau không an toàn chỉ nhờ một ngưỡng cosine.
  19. 10Vector Stores in ProductionWhen exact search is enough and when you need an approximate index, what it costs in recall, and three operational jobs: adding documents, deleting them, and changing the embedding model without silently breaking everything.
  20. 10Kho vector khi vận hànhKhi nào tìm kiếm chính xác là đủ và khi nào cần chỉ mục gần đúng, cái giá về độ phủ, và ba việc vận hành: thêm, xóa tài liệu và đổi mô hình embedding mà không làm hỏng cả hệ.
  21. 11Observability for Serving SystemsWhat to record for every request so you can diagnose later, reading the metrics the serving engine exposes, the signature each kind of incident leaves, and a post-mortem of a server crash from its own logs.
  22. 11Quan trắc hệ phục vụCần ghi gì cho mỗi yêu cầu để chẩn đoán về sau, đọc các đại lượng bộ máy phục vụ tự đo, dấu vết mỗi loại sự cố để lại, và mổ một lần máy chủ sập từ chính log của nó.
  23. 12End-to-End Project: Deploying a RAG SystemDeploying a retrieval-augmented QA system for multiple users, measuring each stage, and reading results that include the metrics that got worse.
  24. 12Project end-to-end: triển khai hệ thống RAGTriển khai một hệ hỏi đáp có truy hồi cho nhiều người dùng, đo từng giai đoạn và đọc cả những chỉ số trở nên kém hơn.
  25. 13Hybrid Search: Why BM25 and Embeddings Need Each OtherA retrieval mental model for combining exact lexical evidence with semantic similarity.
  26. 14The Evaluation Loop Is Part of the ProductWhy AI quality work should connect test cases, traces, failure classes and product changes in one loop.