Computer VisionComputer Vision
How models learn from images, from pixels and CNNs to detection, segmentation, vision transformers, and CLIP.Cách mô hình học từ hình ảnh, từ pixel và CNN đến detection, segmentation, vision transformer và CLIP.
13 articlesbài viết
- 01IoU → NMS: The Small Geometry Behind Object DetectionFrom overlapping boxes to the suppression rule that quietly shapes detector output.→
- 02What an Image Really Is as DataHow an image becomes a tensor, the four numbers that decide everything (shape, dtype, value range and channel order), and a few silent preprocessing mistakes that drag a model down to guessing.→
- 02Ảnh thực chất là dữ liệu gìMột tấm ảnh biến thành tensor thế nào, bốn con số quyết định là hình dạng, kiểu dữ liệu, thang giá trị và thứ tự kênh màu, và vài lỗi tiền xử lý âm thầm kéo mô hình xuống mức đoán bừa.→
- 03From Convolution to LeNet, AlexNet and VGGComputing the output size, parameter count and receptive field of a stack of convolutions, what stride does to all three, and reading three classic architectures as three answers to the same problem.→
- 03Từ tích chập tới LeNet, AlexNet và VGGTính kích thước đầu ra, số tham số và trường tiếp nhận của một chồng tầng tích chập, stride làm gì với ba con số ấy, và đọc ba kiến trúc kinh điển như ba câu trả lời cho cùng một bài toán.→
- 04ResNet and Skip ConnectionsHow skip connections help deep networks train, how to distinguish depth degradation from overfitting, and what a small experiment reproduces and misses.→
- 04ResNet và kết nối tắtCách kết nối tắt giúp huấn luyện mạng sâu, cách phân biệt suy giảm do độ sâu với overfitting, và một thí nghiệm nhỏ tái hiện được gì và bỏ lỡ gì.→
- 05What a CNN Actually LearnsReading feature maps at each depth, implementing Grad-CAM in a dozen lines, and why a plausible-looking heatmap does not prove the model is reasoning correctly.→
- 05Mô hình CNN thực sự học gìĐọc bản đồ đặc trưng ở từng độ sâu, tự cài Grad-CAM trong hơn chục dòng, và vì sao một heatmap trông hợp lý không chứng minh mô hình suy luận đúng.→
- 06Image Datasets and AugmentationHow often the most common augmentation recipe corrupts the label, measured rather than guessed, and a diagnostic rule for when augmentation is worth turning on.→
- 06Dataset ảnh và augmentationCông thức augmentation phổ biến nhất làm sai nhãn bao nhiêu phần trăm số lần, đo bằng số, và một quy tắc chẩn đoán để biết khi nào nên bật augmentation.→
- 07Transfer Learning and Image Classification in PracticeChoosing how much to fine-tune for the data you have instead of unfreezing everything, reading the accuracy, speed and size trade-off to pick a backbone, and why this chapter and Deep Learning give opposite answers.→
- 07Transfer learning và phân loại ảnh trong thực tếChọn mức fine-tune hợp với lượng dữ liệu thay vì mặc định mở hết, đọc bảng đánh đổi accuracy, tốc độ và kích thước để chọn backbone, và vì sao chương này với chương Deep Learning trả lời ngược nhau.→
- 08Object Detection I: From Bounding Boxes to mAPComputing IoU by hand and why the IoU threshold is a choice, running NMS in your head and where it behaves counter-intuitively, then building the precision-recall curve to get AP.→
- 08Phát hiện vật thể I: từ bounding box tới mAPTính IoU bằng tay và vì sao ngưỡng IoU là một lựa chọn, chạy NMS trong đầu và chỗ nó trái trực giác, rồi dựng đường precision-recall để ra con số AP.→
- 09Object Detection II: Architecture FamiliesPlacing any object detector on the two-stage / one-stage and anchor-based / anchor-free tree, then reading the accuracy, speed and size trade-off measured on the same set of images.→
- 09Phát hiện vật thể II: các họ kiến trúcĐặt mọi mô hình phát hiện vật thể vào cây hai giai đoạn / một giai đoạn và anchor-based / anchor-free, rồi đọc bảng đánh đổi độ chính xác, tốc độ, kích thước đo thật trên cùng một tập ảnh.→
- 10Image Segmentation: From Classes to Individual InstancesFour vision tasks on one ladder, telling semantic from instance segmentation with numbers rather than definitions, and reading Dice and IoU for masks, with a warning about the ground truth itself.→
- 10Phân đoạn ảnh: từ lớp tới từng thực thểBốn bài toán thị giác trên cùng một bậc thang, phân biệt phân đoạn ngữ nghĩa với phân đoạn thực thể bằng số liệu, và đọc Dice với IoU cho mặt nạ, kèm cảnh báo về chính nhãn thật.→
- 11Vision TransformerHow vision transformers turn images into patches, which CNN inductive biases they relax, and what changes in a controlled comparison.→
- 11Vision TransformerCách vision transformer chia ảnh thành các patch, những inductive bias nào của CNN được nới lỏng, và điều gì thay đổi trong một phép so sánh có kiểm soát.→
- 12Representation Learning: From Fixed Labels to a Vector SpacePulling embeddings out of a trained network to find similar images, cluster without labels and classify with no training at all using CLIP, and reading those numbers without over-claiming.→
- 12Học biểu diễn: từ nhãn cố định tới không gian vectorLấy vector biểu diễn từ một mạng đã huấn luyện để tìm ảnh giống, gom cụm không cần nhãn, phân loại không cần huấn luyện bằng CLIP, và đọc những con số ấy cho đúng.→
- 13End-to-End Project: Reading Invoices from PhotosAssembling several vision components into a complete system, seeing in numbers that the hardest part usually is not the model, and a checklist for handing the system over.→
- 13Project end-to-end: đọc hóa đơn từ ảnhGhép nhiều thành phần thị giác thành một hệ hoàn chỉnh, thấy bằng số liệu rằng phần khó nhất thường không nằm ở mô hình, và một danh sách kiểm tra để bàn giao hệ.→