Machine LearningMachine Learning

Classic algorithms worked through on real data with scikit-learn, from linear models to trees, clustering, and end-to-end projects.Các thuật toán kinh điển được triển khai trên dữ liệu thật bằng scikit-learn, từ mô hình tuyến tính đến cây, phân cụm và các project end-to-end.

19 articlesbài viết

  1. 01Your First Project: Predicting House Prices with scikit-learnRun a complete machine learning workflow yourself on real house-price data: state the problem and the metric, split train/test, train with a Pipeline, measure RMSE/MAE against a baseline, read the largest errors, then save the model.
  2. 01Project đầu tay: dự đoán giá nhà với scikit-learnTự tay chạy trọn một quy trình machine learning trên dữ liệu giá nhà thật: nêu bài toán và thước đo, chia train/test, huấn luyện bằng Pipeline, đo RMSE/MAE so với baseline, đọc lỗi lớn nhất rồi lưu mô hình.
  3. 02What Does It Mean for a Machine to Learn?A clean model of supervised learning: data, hypothesis, objective, optimization and generalization.
  4. 03Linear Regression: The Smallest Useful Learning ModelUse linear regression to understand features, parameters, residuals, loss and what a model is really fitting.
  5. 04Linear Regression: What the Line Tells YouRead regression coefficients the right way, understand why least squares is exactly MSE, use R² to see how well the line fits, spot the three signs that the linearity assumption is failing, and learn how to extend the model.
  6. 04Hồi quy tuyến tính: đường thẳng nói gì với bạnĐọc hệ số hồi quy đúng cách, hiểu vì sao bình phương tối thiểu chính là MSE, dùng R² để biết đường thẳng khớp đến đâu, nhận ra ba dấu hiệu giả định tuyến tính đang sai và cách mở rộng mô hình.
  7. 05Logistic Regression: From Scores to Class ProbabilitiesWhy a linear decision boundary plus a sigmoid becomes a powerful baseline for binary classification.
  8. 06Logistic Regression: Classification with ProbabilitiesWhy a straight line is wrong for 0/1 labels, how the sigmoid turns a score into a probability, how to read coefficients through odds, why log loss is exactly cross-entropy, and how to run logistic regression inside a Pipeline on the adult dataset.
  9. 06Hồi quy logistic: phân loại bằng xác suấtVì sao không dùng đường thẳng cho nhãn 0/1, sigmoid biến điểm số thành xác suất ra sao, đọc hệ số qua odds, log loss chính là cross-entropy, và chạy hồi quy logistic trong một Pipeline trên bộ dữ liệu adult.
  10. 07Decision Trees: Learning by Splitting the Feature SpaceUnderstand recursive splits, impurity, overfitting and why tree ensembles are so strong on structured data.
  11. 08Regularization: Ridge, Lasso and the Art of Restraining a ModelRestrain a model with too many features by penalizing large coefficients: Ridge shrinks evenly, Lasso cuts all the way to 0, ElasticNet blends the two, why you must scale before penalizing, and how to choose alpha with cross-validation.
  12. 08Regularization: Ridge, Lasso và nghệ thuật kìm mô hìnhKìm mô hình có quá nhiều feature bằng khoản phạt hệ số lớn: Ridge co đều, Lasso cắt hẳn về 0, ElasticNet lai hai kiểu, vì sao phải scale trước khi phạt và cách chọn alpha bằng cross-validation.
  13. 09Generalization, Overfitting and RegularizationWhy fitting training data is not the goal, how overfitting appears and what regularization is really trying to control.
  14. 10k-Nearest Neighbors: Learning from the NeighborsWhy k-NN learns nothing at training time yet still makes predictions, computing a prediction by hand, choosing k with cross-validation, why features must be scaled, the curse of dimensionality, and when k-NN is the right choice.
  15. 10k-Nearest Neighbors: học từ hàng xómVì sao k-NN không học gì lúc train mà vẫn dự đoán được, tự tính tay một dự đoán, chọn k bằng cross-validation, vì sao phải scale feature, lời nguyền số chiều và khi nào nên dùng k-NN.
  16. 11Decision Trees: A Model You Can Read Like a FlowchartHow a decision tree picks its questions using Gini purity, reading a real tree trained on the adult data, why deep trees overfit and how to rein them in, how regression trees predict in steps, and the pros and cons of trees.
  17. 11Cây quyết định: mô hình đọc được như một sơ đồCây quyết định chọn câu hỏi bằng độ sạch Gini ra sao, đọc một cây thật huấn luyện trên dữ liệu adult, vì sao cây sâu học tủ và cách kìm nó, cây hồi quy dự đoán kiểu bậc thang và ưu nhược của cây.
  18. 12Model Evaluation: Measure the Failure You Actually Care AboutAccuracy, precision, recall, F1, ROC, regression error and the deeper problem of matching metrics to decisions.
  19. 13Random Forest & Gradient Boosting: The Power of the CrowdWhy combining hundreds of unstable trees yields a stable model: bagging and random forest grow trees in parallel, boosting grows them in sequence, how to read out-of-bag and feature importance, and why tree models tend to win on tabular data.
  20. 13Random Forest & Gradient Boosting: sức mạnh của đám đôngVì sao gom hàng trăm cây bất ổn lại được một mô hình ổn định: bagging và random forest trồng song song, boosting trồng nối tiếp, đọc out-of-bag và feature importance, và lý do mô hình cây hay thắng trên dữ liệu bảng.
  21. 14SVM: Finding the Widest MarginAmong the countless lines that separate two classes, SVM picks the one with the widest margin: support vectors, the C knob for the soft margin, the kernel trick and the gamma of the RBF kernel, running SVC on breast_cancer, and when SVM is worth using.
  22. 14SVM: tìm lề phân cách rộng nhấtTrong vô số đường tách được hai lớp, SVM chọn đường có lề rộng nhất: support vector, núm C cho lề mềm, kernel trick và gamma của kernel RBF, chạy SVC trên breast_cancer và khi nào SVM đáng dùng.
  23. 15Clustering with k-means: Grouping Without LabelsHow k-means differs from classification, how to choose the number of clusters with the elbow method and silhouette score, and when it fails.
  24. 15Phân cụm với k-means: gom nhóm khi không có nhãnK-means khác phân loại ở đâu, cách chọn số cụm bằng elbow method và silhouette score, cùng những trường hợp thuật toán thất bại.
  25. 16PCA: Reducing Dimensions and Choosing How Many Components to KeepWhich direction PCA picks for its new axes and why, how to read the explained variance ratio chart to choose the number of components, how to use PCA both to plot high-dimensional data and as a preprocessing step, and its three limitations.
  26. 16PCA: giảm chiều và chọn số thành phần cần giữPCA tìm trục mới theo hướng nào và vì sao, đọc biểu đồ tỉ lệ phương sai giải thích để chọn số thành phần, dùng PCA để vẽ dữ liệu nhiều chiều và làm tiền xử lý, cùng ba giới hạn của nó.
  27. 17Feature Engineering & Pipelines: Making Data SpeakThe most common ways to create new features and why they help a model, choosing an encoding for text columns, the most subtle kind of leakage in preprocessing, and wrapping everything into a Pipeline with ColumnTransformer.
  28. 17Feature engineering & Pipeline: làm dữ liệu biết nóiNhững cách tạo đặc trưng mới thường dùng và vì sao chúng giúp mô hình, chọn cách mã hóa cho cột chữ, kiểu lộ đề tinh vi nhất trong tiền xử lý, và gói toàn bộ vào Pipeline với ColumnTransformer.
  29. 18End-to-End Project: Predicting Income from a ProfileRun a complete machine learning project on the adult data: pick a metric for a class-imbalanced problem, build a pipeline, compare models with cross-validation, choose a threshold by cost, explain the model, check for gaps between groups, and hand it over.
  30. 18Project end-to-end: dự đoán thu nhập từ hồ sơChạy trọn một dự án học máy trên dữ liệu adult: chọn metric cho bài toán lệch lớp, dựng pipeline, so mô hình bằng cross-validation, chọn ngưỡng theo chi phí, giải thích mô hình, kiểm tra chênh lệch giữa các nhóm và bàn giao.
  31. 19Linear Regression From ScratchImplement the full learning loop with arrays, gradients and no machine-learning framework.