Async vs Parallelism: Why Async May Not Make Your Backend Faster
Understand concurrency and parallelism, event loops, blocking code, thread and process pools, workers, and how to benchmark async against a real workload.
Blog
Deep dives into real-world AI systems, papers, experiments, and engineering lessons beyond the handbook's core learning path.Các bài đào sâu về hệ thống AI thực tế, paper, thử nghiệm và bài học engineering nằm ngoài lộ trình cốt lõi của handbook.
Understand concurrency and parallelism, event loops, blocking code, thread and process pools, workers, and how to benchmark async against a real workload.
Phân biệt concurrency và parallelism, hiểu event loop, blocking code, thread/process pools, workers và cách benchmark async đúng với workload thực tế.
Fixed and sliding windows, token buckets, atomic Redis checks, HTTP 429, concurrency limits, and testing a rate limiter across replicas.
Từ fixed window, sliding window và token bucket đến Redis Lua, HTTP 429, giới hạn concurrency và cách kiểm thử rate limiter trên nhiều replica.
From cache-aside, TTLs, and invalidation to read-write races, cache stampedes, and measuring whether Redis makes an API faster without making it wrong.
Từ cache-aside, TTL và invalidation đến race condition, cache stampede và cách đo xem Redis có thực sự giúp API nhanh mà vẫn đúng hay không.
Why fast SQL can sit behind slow connection checkout, how SQLAlchemy pools scale across workers, and when PgBouncer helps or merely moves the queue.
SQL chạy nhanh nhưng API vẫn chậm: tìm thời gian chờ connection, tính pool trên toàn bộ workers và hiểu khi nào PgBouncer thực sự có ích.
How fact grain, dimensions, SCD Type 2, conformed dimensions, and analytical marts prevent double counting and make business metrics trustworthy.
Cách định nghĩa grain, thiết kế Fact và Dimension, quản lý SCD Type 2 và Data Mart để tránh tính trùng và giữ các chỉ số kinh doanh đáng tin cậy.
How to test data completeness, uniqueness, business rules, freshness, reconciliation, and publication before a successful pipeline produces a misleading dashboard.
Kiểm tra tính đầy đủ, trùng lặp, quy tắc nghiệp vụ, độ mới và đối soát dữ liệu trước khi một pipeline thành công tạo ra dashboard sai.
How PostgreSQL WAL, logical decoding, Debezium, snapshots, and replication slots feed analytics, and why reliable CDC still needs ordering, deduplication, and reconciliation.
Cách PostgreSQL WAL, Logical Decoding, Debezium, Snapshot và Replication Slot đưa thay đổi sang Analytics, cùng những bài toán về thứ tự, dữ liệu trùng và đối soát.
Why analytical queries can slow transactional APIs, and how data models, materialized views, replicas, pipelines, and warehouses fit different workloads.
Vì sao dashboard có thể làm chậm API giao dịch, và cách chọn giữa tối ưu SQL, Materialized View, Read Replica và Data Warehouse theo workload thực tế.
A practical introduction to PostgreSQL transactions, MVCC, isolation levels, locking, optimistic concurrency, and the limits of ACID across services.
Tìm hiểu Transactions, MVCC, isolation levels và các cách xử lý race condition khi nhiều request cùng cập nhật dữ liệu trong PostgreSQL.
A practical guide to PostgreSQL query plans, B-tree and composite indexes, selectivity, EXPLAIN ANALYZE, and the costs of indexing a real workload.
Vì sao PostgreSQL có thể không dùng Index vừa tạo, cách đọc Query Plan, thiết kế Composite Index và đánh giá chi phí trên toàn bộ workload.
How prompt injection crosses the boundary between data and instructions, and how to protect Agent actions with least privilege, authorization, and evaluation.
Prompt injection khai thác ranh giới giữa dữ liệu và chỉ thị như thế nào, và cách giới hạn quyền, kiểm tra hành động, đánh giá phòng vệ trong AI Agent.
Designing durable Agent execution with checkpoints, human approval, retries, idempotency, and a state machine that prevents duplicate actions.
Thiết kế Agent chạy dài với checkpoint, human approval, retry, idempotency và state machine để phục hồi mà không nhân đôi hành động.
A practical guide to choosing tools, Agent Skills, MCP, subagents, or ordinary workflows - and knowing where each responsibility belongs.
Phân biệt vai trò của Tool, Agent Skill, MCP, Subagent và Workflow để mở rộng khả năng AI Agent đúng chỗ, đúng quyền.
How to build reliable Agent memory in production, from preferences and retrieval to updates, conflicts, deletion, and evaluation.
Từ ghi nhớ sở thích đến cập nhật, truy xuất và xóa đúng lúc: cách thiết kế Agent Memory đáng tin cậy trong production.
A practical look inside an Agent's Context Builder: authorized retrieval, evidence selection, tool-output shaping, memory, handoffs, and evaluation.
Đi sâu vào Context Builder của AI Agent: kiểm soát quyền truy cập, chọn bằng chứng, xử lý kết quả tool, quản lý context qua nhiều bước và đánh giá pipeline.
A practical architecture for running golden test cases, capturing traces, combining specialized graders, diagnosing failures, and comparing AI system versions.
Cách chạy Golden Dataset, thu thập traces, kết hợp các grader chuyên trách, chẩn đoán lỗi và so sánh các phiên bản AI System.
How to evaluate answers, retrieval, tool use, final outcomes, and reliability with code-based graders, references, LLM judges, and human review.
Cách đánh giá câu trả lời, retrieval, tool use, kết quả cuối và độ ổn định bằng code-based grader, reference, LLM Judge và human review.
A practical guide to finding the source of truth, generating test cases, validating expected outcomes, and maintaining a golden dataset for chatbots, RAG, and AI Agents.
Từ nguồn sự thật, cách tạo test case đến xác minh kết quả kỳ vọng và duy trì Golden Dataset cho chatbot, RAG và AI Agent.
From feature matching and Structure-from-Motion to meshes, NeRF, and 3D Gaussian Splatting: a practical guide to recovering 3D scenes from overlapping photos and knowing what the result can actually do.
Từ feature matching, Structure-from-Motion tới mesh, NeRF và 3D Gaussian Splatting: cách khôi phục scene 3D từ ảnh chụp nhiều góc và hiểu kết quả dùng được vào việc gì.
A factory's 3D model can show where machines are. A digital twin connects assets, operational data, and context to help people understand what is happening and evaluate what might happen next.
Mô hình 3D cho thấy máy móc nằm ở đâu. Digital Twin còn kết nối thiết bị, dữ liệu vận hành và ngữ cảnh để hiểu điều gì đang diễn ra và cân nhắc điều gì có thể xảy ra tiếp theo.
When expert labels are scarce, active learning helps choose images worth reviewing. Understand uncertainty, diversity, random baselines, annotation quality, and the human-in-the-loop workflow.
Khi nhãn từ chuyên gia có hạn, Active Learning giúp chọn ảnh đáng xem tiếp. Tìm hiểu uncertainty, diversity, random baseline, chất lượng nhãn và workflow có con người trong vòng lặp.
How can millions of unlabeled images become useful training data? A practical introduction to self-supervised learning through SimCLR, MAE, DINO, and downstream evaluation.
Làm sao biến hàng triệu ảnh chưa gán nhãn thành dữ liệu học hữu ích? Tìm hiểu self-supervised learning qua SimCLR, MAE, DINO và cách đánh giá trên task thực tế.
Synthetic data lets teams create rare scenes, controlled variations, and labels for computer vision and robotics. Learn where it helps, why sim-to-real gaps remain, and how to test it on real data.
Synthetic data giúp chủ động tạo tình huống hiếm, biến thể có kiểm soát và nhãn cho computer vision, robotics. Tìm hiểu lợi ích, khoảng cách sim-to-real và cách đánh giá bằng dữ liệu thật.
A model may perform beautifully on a test set yet struggle when cameras, lighting, products, or time change. Learn how to test across domains, spot shortcuts, monitor drift, and respond.
Model chạy rất tốt trên test set nhưng dễ sai khi camera, ánh sáng, sản phẩm hoặc thời gian thay đổi. Tìm hiểu cách test theo domain, phát hiện shortcut, theo dõi drift và xử lý.
Deploying AI near a camera or robot means balancing quality with end-to-end latency, runtime memory, power, heat, and device operations. A practical guide to the trade-offs behind Edge AI.
Đưa AI đến gần camera hay robot là bài toán cân bằng chất lượng với latency toàn pipeline, bộ nhớ lúc chạy, điện năng, nhiệt và vận hành thiết bị. Một cách nhìn thực tế về Edge AI.
A VLM answers questions about visual input; a VLA produces robot actions. Understand action tokens, robot data, proprioception, closed-loop control, and why safety changes the problem.
VLM trả lời câu hỏi về hình ảnh; VLA tạo hành động cho robot. Tìm hiểu action token, dữ liệu robot, proprioception, vòng lặp quan sát-hành động và giới hạn an toàn của hai cách tiếp cận.
Physical AI goes beyond recognizing a scene: a system must sense, plan, act, and check what actually happened. A practical guide to sensors, world models, simulation, latency, and safety.
Physical AI không chỉ nhận ra cảnh vật: cả hệ thống phải cảm nhận, lập kế hoạch, hành động và kiểm tra kết quả. Tìm hiểu cảm biến, world model, mô phỏng, độ trễ và an toàn.
From detection and segmentation to VLMs, VLAs, and Physical AI: language makes visual systems more flexible, but safe action still needs perception, planning, control, and feedback.
Từ detector và segmentation đến VLM, VLA và Physical AI: ngôn ngữ mở rộng khả năng đặt câu hỏi về hình ảnh, nhưng hành động an toàn vẫn cần perception, planning, control và feedback.
A case study of Uber's published choices: an OpenAI-compatible API, multiple model backends, PII redaction, security review, auditing, cost controls, and the trade-offs of a shared gateway.
Mổ xẻ những lựa chọn Uber công khai: API tương thích OpenAI, nhiều nguồn model, PII redaction, security review, audit, cost và các trade-off của một gateway tập trung.
A practical guide to defining success, building task-based evals, grading RAG and Agent behavior, comparing model changes, and learning from production failures.
Cách xác định thành công, xây eval theo task, chấm RAG và Agent, so sánh thay đổi model và học từ những lỗi xảy ra trong production.
Turn production feedback and traces into root-cause analysis, regression cases, targeted fixes, and safer rollouts.
Từ feedback và trace trong production đến điều tra lỗi, thêm regression case, sửa đúng thành phần và kiểm chứng trước khi rollout.
Why HTTP 200 is not enough for an AI product: logs, metrics, traces, RAG evidence, Agent steps, asynchronous work, token usage, privacy, and evaluation.
Vì sao HTTP 200 chưa đủ để đánh giá sản phẩm AI: logs, metrics, traces, bằng chứng RAG, hành động của Agent, xử lý bất đồng bộ, token, quyền riêng tư và evaluation.
Why an AI system may publish events instead of chaining service calls, and how Kafka topics, consumer groups, partitions, replay, and failure handling fit into the picture.
Vì sao hệ thống AI có thể phát event thay vì gọi service nối tiếp, và Kafka topic, consumer group, partition, replay cùng xử lý lỗi nằm ở đâu trong bức tranh đó?
How an LLM reuses attention state while generating, why long contexts and concurrent requests consume GPU memory, and how paged allocation and prefix caching help.
Model giữ lại state của attention khi sinh token như thế nào, vì sao context dài và nhiều request cùng chạy tốn GPU memory, cùng vai trò của PagedAttention và prefix caching.
What happens between thousands of incoming prompts and a finite pool of GPUs? An introduction to queues, schedulers, continuous batching, prefill, decode, and latency.
Giữa hàng nghìn prompt và số GPU có hạn là những gì? Bài viết giải thích queue, scheduler, continuous batching, prefill, decode và các thước đo latency.
An accessible guide to the control layer between AI applications and models: credentials, quotas, routing, fallback, observability, policy, caching, and the trade-offs of centralization.
Một cách hiểu dễ tiếp cận về lớp kiểm soát giữa ứng dụng AI và model: API key, quota, routing, fallback, observability, policy, cache và cái giá của việc tập trung hóa.
When should one AI Agent delegate work? A practical guide to specialist Agents, orchestration, handoffs, parallel work, shared state, security, cost, and evaluation.
Khi nào nên chia việc cho nhiều AI Agent? Bài viết giải thích specialist, orchestration, handoff, xử lý song song, shared state, bảo mật, chi phí và evaluation.
What changes when a model steers a workflow: agent loops, state, tools, runtime controls, approval, stopping rules, tracing, and when one Agent is enough.
Điều gì thay đổi khi model tham gia điều khiển workflow? Bài viết giải thích agent loop, state, tools, runtime, approval, điều kiện dừng, tracing và khi nào một Agent là đủ.
Follow evidence through a real RAG pipeline: ingestion, parsing, chunking, indexing, hybrid search, permissions, reranking, context building, generation, and evaluation.
Theo hành trình của evidence qua ingestion, parsing, chunking, indexing, hybrid search, kiểm soát quyền, reranking, context building, generation và evaluation.
Why AI needs retrieval for current and private knowledge, how RAG works, when it beats a long prompt or fine-tuning, and why retrieval quality matters.
RAG bắt đầu từ vấn đề thiếu thông tin, không phải thiếu trí thông minh. Bài viết giải thích khi nào cần retrieval, RAG hoạt động ra sao và vì sao lấy đúng nguồn rất quan trọng.
Follow one chat request from Enter to the first visible text: identity, conversation state, context, tools, model calls, streaming, tracing, and failure handling.
Theo chân một message từ lúc nhấn Enter đến khi thấy chữ đầu tiên: xác thực, trạng thái hội thoại, context, tools, model, streaming, tracing và xử lý lỗi.
A practical map of GenAI architecture: start with a model call, then add context, retrieval, memory, tools, orchestration, gateways, security, and evals only when needed.
Bản đồ dễ hiểu về kiến trúc GenAI: bắt đầu từ một lần gọi model, rồi chỉ thêm context, RAG, memory, tools, orchestration, gateway và eval khi cần.
A compelling AI demo proves an idea can work. Production also demands reliable quality, speed, cost, security, and recovery under real traffic.
Một demo ấn tượng mới chứng minh ý tưởng có thể chạy. Production còn đòi hỏi chất lượng, tốc độ, chi phí, bảo mật và khả năng phục hồi với người dùng thật.
A high benchmark score measures success on a defined test. Real work also demands judgment, reliable workflows, and evaluations that reflect your users.
Điểm benchmark rất hữu ích, nhưng không nói hết AI sẽ làm được gì trong sản phẩm của bạn. Bài viết giải thích khoảng cách ấy và cách xây eval sát công việc thực tế.
Why real AI systems may combine a general model with smaller or specialized models, how model routing works, and when the extra complexity is actually worth it.
Không phải task nào cũng cần model mạnh nhất. Bài viết giải thích cách multi-model system và model routing kết hợp model tổng quát với những model nhỏ hoặc chuyên biệt, cùng những đánh đổi đi kèm.
A practical introduction to world models, how they predict the consequences of actions, how V-JEPA 2 and Genie 3 approach the problem, and why simulation is useful but never perfect.
World Model giúp AI dự đoán điều gì có thể xảy ra sau một hành động. Bài viết giải thích cách V-JEPA 2 và Genie 3 tiếp cận bài toán này, vì sao mô phỏng hữu ích và tại sao nó không bao giờ là bản sao hoàn hảo của thế giới.
A non-technical guide to context windows, short-term and long-term memory, how memory differs from RAG, and why a useful AI system must know both what to remember and what to forget.
Vì sao AI có thể rất thông minh nhưng vẫn hay quên? Bài viết giải thích cách context và memory hoạt động, memory khác RAG ra sao, và vì sao một hệ thống AI tốt phải biết điều gì nên nhớ, điều gì nên quên.
A practical guide to how MCP connects agents to tools, how A2A connects independent agents, and why real systems may need both.
Giải thích thực tế về cách MCP kết nối Agent với công cụ, A2A kết nối các Agent độc lập và vì sao một hệ thống có thể cần cả hai.
A non-technical introduction to the Model Context Protocol, how it connects AI applications to tools and data, and what it does not solve.
MCP đang dần trở thành một trong những mảnh ghép quan trọng của hệ sinh thái AI Agent. Nếu Agent muốn đọc dữ liệu, gọi công cụ hay làm việc với những hệ thống bên ngoài, nó cần một cách kết nối đủ thống nhất - và MCP được tạo ra để giải quyết chính bài toán đó.
A practical way to decide when AI should respond quickly, when it should reason deeply, and why production systems often need both.
Một cách thực tế để xác định khi nào AI nên phản hồi nhanh, khi nào cần suy luận sâu và vì sao hệ thống production thường cần cả hai.
Jev is TypeSafe AI's first System One Model, designed to return structured decisions that software can use directly.
Jev là System One Model đầu tiên của TypeSafe AI, được thiết kế để đưa ra các quyết định có cấu trúc mà phần mềm có thể sử dụng trực tiếp.