Generative AIGenerative AI
How generative models work, from tokenization and attention to prompting, retrieval, diffusion, and RAG.Cách các mô hình sinh hoạt động, từ tokenization và attention đến prompting, retrieval, diffusion và RAG.
22 articlesbài viết
- 01Tokenization: How Text Becomes Model InputWhy language models operate on tokens rather than words, how subword tokenization works and what token boundaries affect.→
- 01Tokenization: văn bản trở thành đầu vào của model như thế nào?Vì sao model ngôn ngữ xử lý token thay vì từ, subword tokenization hoạt động ra sao và ranh giới token ảnh hưởng những gì.→
- 02What a Language Model Really DoesA language model is trained to do one very small thing, why repeating it produces text, and a transformer trained on this very series going from random characters to Vietnamese sentences.→
- 02Mô hình ngôn ngữ thực chất làm gìMô hình ngôn ngữ được huấn luyện để làm đúng một việc rất nhỏ, vì sao lặp lại việc ấy sinh ra văn bản, và một transformer tự huấn luyện trên chính loạt bài này đi từ ký tự ngẫu nhiên tới câu tiếng Việt.→
- 03Attention: Let Each Token Decide What MattersA geometric and operational explanation of queries, keys, values and why attention changed sequence modeling.→
- 03Attention: để mỗi token chọn thông tin cần chú ýGiải thích query, key, value theo góc nhìn hình học và vận hành, cùng lý do attention thay đổi cách xử lý chuỗi.→
- 04Tokenizers and the Cost of VietnameseHow tokenizers split text, preserve information, and why Vietnamese text may require more tokens than equivalent English text.→
- 04Tokenizer và cái giá của tiếng ViệtCách tokenizer chia nhỏ văn bản mà vẫn giữ nguyên thông tin, và vì sao tiếng Việt có thể cần nhiều token hơn nội dung tiếng Anh tương đương.→
- 05Attention: Queries, Keys, ValuesThe mechanism the previous two chapters deliberately left out, computing attention by hand, and why an attention matrix must not be read as an explanation of the model's answer.→
- 05Attention: truy vấn, khóa, giá trịCơ chế mà hai chương trước cố ý để lại, tính một phép attention bằng tay, và vì sao không được đọc ma trận attention như lời giải thích cho câu trả lời của mô hình.→
- 06Transformers: A Stack for Building Contextual RepresentationsHow embeddings, attention, feed-forward layers, residual paths and normalization combine into the transformer architecture.→
- 06Transformer: xếp nhiều lớp để tạo biểu diễn có ngữ cảnhEmbedding, attention, feed-forward, đường residual và normalization kết hợp như thế nào trong kiến trúc transformer.→
- 07LLM Training vs Inference: Two Very Different SystemsPretraining, instruction tuning and next-token generation explained as separate stages with different engineering constraints.→
- 07Huấn luyện và inference LLM: hai bài toán hệ thống rất khác nhauPhân biệt pretraining, instruction tuning và quá trình sinh từng token, cùng những ràng buộc kỹ thuật khác nhau của từng giai đoạn.→
- 08The Transformer from ScratchAssembling a full Transformer from separate building blocks, the problem each block solves, and a table measuring how much breaks when each one is removed, including one counter-intuitive result.→
- 08Transformer từ đầuRáp một Transformer đầy đủ từ những viên gạch rời, mỗi viên gạch giải quyết vấn đề gì, và bảng đo tháo từng viên ra thì hỏng đến đâu, có cả một kết quả ngược trực giác.→
- 09RAG From First Principles: Retrieve Evidence, Then GenerateA clean architecture for retrieval-augmented generation, what retrieval actually contributes and the failure modes hidden by simple diagrams.→
- 09RAG từ nguyên lý đầu tiên: tìm bằng chứng rồi mới sinh câu trả lờiKiến trúc RAG rõ ràng, vai trò thực sự của retrieval và những kiểu lỗi thường bị che khuất bởi sơ đồ đơn giản.→
- 10Text Generation: The KnobsThe decoding step between a probability distribution and actual text, what temperature, top-k and top-p do to that distribution, and choosing settings from measurements rather than feel.→
- 10Sinh văn bản: các núm vặnTừ phân bố xác suất tới câu chữ còn một bước quyết định, temperature, top-k và top-p làm gì với phân bố, và chọn cấu hình theo số đo thay vì cảm giác.→
- 11Context Engineering: Decide What the Model Gets to SeePrompts are only one part of the problem. Context engineering manages instructions, state, retrieval, tools and memory under a finite budget.→
- 11Context Engineering: quyết định model được nhìn thấy gìPrompt chỉ là một phần. Context Engineering quản lý chỉ dẫn, state, retrieval, tools và memory trong một ngân sách context hữu hạn.→
- 12Model Scale: What We Know and What We Don'tWhere a small model breaks, why you cannot infer scaling laws from one or two models, and why parameter count is not the only thing that decides capability.→
- 12Quy mô mô hình: cái gì ta biết và cái gì khôngMột mô hình nhỏ gãy ở đâu, vì sao không thể suy từ một hai mô hình ra quy luật về quy mô, và vì sao số tham số không phải biến duy nhất quyết định năng lực.→
- 13From Base Model to AssistantHow a model behaves before post-training, the order and names of the training stages, and why the whole change in behavior should not be credited to a single technique.→
- 13Từ mô hình gốc tới trợ lýMột mô hình chưa qua post-training hành xử ra sao, thứ tự và tên gọi các giai đoạn huấn luyện, và vì sao không quy toàn bộ khác biệt hành vi cho riêng một kỹ thuật.→
- 14Tool Calling and Agents: Give the Model Actions, Not OmnipotenceA grounded model of tools, agent loops, state and why reliable control flow matters more than calling everything an agent.→
- 14Tool Calling và Agent: cho model khả năng hành động, không phải toàn quyềnCách nhìn thực tế về tools, vòng lặp Agent, state và lý do luồng điều khiển đáng tin quan trọng hơn việc gọi mọi thứ là Agent.→
- 15Evaluating Generative AI: From Vibes to Failure TaxonomiesHow to evaluate RAG, tool use and generated answers with layered metrics, golden cases, traces and targeted human review.→
- 15Đánh giá Generative AI: từ cảm giác đến phân loại lỗiCách đánh giá RAG, tool use và câu trả lời được tạo bằng nhiều lớp metric, bộ ca chuẩn, trace và kiểm tra có chọn lọc của con người.→
- 16Prompts: What Actually Changes the ResultMeasuring what a prompt does against gradable answers instead of gut feeling, a result that contradicts popular advice, and where prompting stops helping.→
- 16Prompt: cái gì thật sự đổi kết quảĐo tác dụng của prompt bằng đáp án chấm được thay vì cảm giác, một kết quả ngược lời khuyên phổ biến, và chỗ prompt hết tác dụng.→
- 17Hallucination and ReliabilityFour different ways a language model gets answers wrong, why its training objective does not guarantee correctness, and measurements showing why "are you sure?" does not work as a filter.→
- 17Ảo giác và độ tin cậyBốn kiểu trả lời sai của mô hình ngôn ngữ, vì sao hàm mục tiêu không bảo đảm tính đúng, và số đo cho thấy vì sao câu hỏi "bạn có chắc không" không dùng làm bộ lọc được.→
- 18Text Embeddings and RetrievalTwo different things both called embeddings, building semantic search over Vietnamese documents, and the trap of a plain topk that always returns k results even when nothing relevant exists.→
- 18Embedding văn bản và truy hồiHai thứ cùng tên embedding nhưng dùng vào hai việc khác nhau, dựng tìm kiếm ngữ nghĩa trên tài liệu tiếng Việt, và cái bẫy topk luôn trả về k kết quả kể cả khi kho không có câu trả lời.→
- 19Diffusion Models: Generating Images from NoiseHow diffusion works, adding noise and learning to remove it, through a model trained from scratch, and where real image generators differ from that toy model.→
- 19Mô hình khuếch tán: sinh ảnh từ nhiễuCơ chế khuếch tán, thêm nhiễu rồi học cách gỡ, qua một mô hình tự huấn luyện từ số 0, và các hệ sinh ảnh thật khác mô hình đồ chơi ấy ở đâu.→
- 20End-to-End Project: Question Answering over This BookBuilding a complete RAG system over the 51 lessons of the previous four chapters, and evaluating retrieval and answering separately so you know exactly where it goes wrong.→
- 20Project end-to-end: hỏi đáp trên chính bộ sách nàyDựng một hệ RAG hoàn chỉnh trên chính 51 bài của bốn chương trước, và đánh giá riêng hai tầng truy hồi và trả lời để biết chính xác hệ sai ở đâu.→
- 21Agent Memory Is a Context Architecture ProblemWhy durable memory is less about storing messages and more about deciding what becomes context, when, and for whom.→
- 21Bộ nhớ của Agent là bài toán kiến trúc contextBộ nhớ lâu dài không chỉ là lưu tin nhắn, mà còn là quyết định thông tin nào được đưa lại vào context, khi nào và cho ai.→
- 22Repository Anatomy: A Minimal Agent Graph RuntimeThe core abstractions hiding underneath graph-based agent frameworks, rebuilt as a small execution model.→
- 22Giải phẫu repository: một graph runtime tối giản cho AgentNhững khái niệm nền tảng phía dưới các framework Agent dạng graph, được dựng lại thành một mô hình thực thi nhỏ.→