Suppose we have records of customers and know which of them repaid a loan. A new customer arrives. No one has written a rule for this particular person, yet we still need a prediction. Machine learning finds patterns in known data and uses them to make predictions. The real question comes after the first correct answer: will the rule work for another customer it has never seen?
We will follow a classification problem from start to finish. First, we will turn objects into features and compare two ways to separate classes. Then we will keep the data used for learning apart from the data used for selection and final evaluation. That distinction explains why flawless performance on familiar examples can be misleading.
What a model learns from data
Patterns instead of a ready-made rule
Machine learning looks for patterns in data and applies them to new objects. A model describes a possible pattern, statistics helps us study it, and an optimization algorithm adjusts the model's parameters using examples. The model is not handed an answer for every future case. It must carry what it learned beyond the list of familiar objects.
Why volume cannot replace quality
With too little data, or data of poor quality, a model cannot learn a useful pattern reliably. A large amount of good data makes it possible to train more flexible models. So before debating a clever algorithm, we should look at what it is given to learn from. Even a diligent student cannot reconstruct the missing pages of a workbook from its handsome cover.
Classification: objects, features, and labels
Credit risk as a two-class problem
Classification starts with objects whose class labels are known. A model learns from them to assign a label to a new object. In the credit-risk example, each customer has two features: annual income and credit-card balance. The known customers fall into two classes: those who repaid a loan and those who defaulted. For a new customer, the class label is unknown. Predicting it is the task.
Let $x$ stand for a customer's features and $\hat y$ for the predicted label. A classifier is a model $h$ that maps one to the other:
The equation is short, but it does not tell us how the model makes its decision. That depends on the chosen classifier and on the features we give it.
When an object is made of pixels
The same language works for images. In the handwritten-digit example, an image of size$m \times n$ can be flattened into a vector $x \in \mathbb{R}^{mn}$, with one coordinate per pixel. Writing out the rows of a $4 \times 4$ image in sequence gives a vector with 16 coordinates. Distinguishing handwritten ones from sevens is then another classification problem, only the feature space is no longer two-dimensional.
How a classifier draws a boundary
With two features, we can picture each customer as a point on a plane. A decision boundary separates regions where the classifier predicts different classes. Different models can divide the same labeled data in different ways. Choosing a model also means choosing an assumption about the shape of the pattern.
A linear classifier
A linear classifier uses a hyperplane. If $w$ represents feature weights and$b$ an offset, points on the boundary satisfy:
In a two-feature space, that hyperplane is a line. The boundary is smooth and does not bend around every training point, but it may be too rigid when the classes have a complicated arrangement. It need not get every known customer's label right.
The nearest-neighbors method
The $k$-nearest-neighbors method takes a different route: for a new object, it finds the$k$ closest objects in the training set and predicts the majority label. With$k = 1$, a single nearest neighbor decides. The method does not calculate an explicit boundary in advance. Each new prediction depends on distances to the stored examples.
A small $k$ produces a complex boundary sensitive to individual points. A large$k$ smooths the boundary but may hide genuine local differences between classes. Which details should we treat as a pattern, and which are accidents of the sample?
Three datasets for three different decisions
To answer the question about new data, we cannot keep showing the model the same examples. Labeled data has different roles: one part adjusts the model, another helps select a candidate, and a third gives a final evaluation of that choice.
The training set
The training set is the part of labeled data on which a learning algorithm adjusts model parameters. For a linear classifier, these include the weights and offset. The phrase “training data” is sometimes used narrowly for this set and sometimes broadly for all the original labeled data. When discussing model quality, naming the specific set is safer than relying on context.
The validation set and hyperparameters
The validation set is held out from parameter training. It lets us compare trained models and choose hyperparameters: settings of the learning algorithm rather than parameters the model learns from the training set. In nearest neighbors, $k$ is a hyperparameter. Candidates are trained on the same training set and compared on the same validation set. The lecture uses an 80% training and 20% validation split as an example, not a required ratio.
The test set
Once a model has been selected, the test set provides a final evaluation on new data, if such a set is available. It is used neither to adjust model parameters nor to choose hyperparameters. An external evaluator may even hide the correct test labels and report only the score. Otherwise, the “final check” soon becomes another chance to fit the answer.
How to measure classification error
The share of wrong answers
For a set $S$, classification error is the fraction of objects to which the classifier assigns the wrong label. If $y_i$ is the true label of object $x_i$, we can write:
The indicator $\mathbb{1}$ is one for a wrong answer and zero for a correct one. The sum counts mistakes, and division by $|S|$ turns that count into a fraction. Change $S$, and the meaning of the result changes even though the equation stays the same.
Training, validation, and test errors
Training error tells us how the model responds to familiar objects. Validation error helps compare candidates and choose hyperparameters. Test error evaluates the selected model on new objects. Validation error is usually higher than training error because the model did not see those examples while its parameters were adjusted. Yet a number without the name of its dataset tells us less about quality than it seems.
With one nearest neighbor, training error is normally zero: each training point is its own closest neighbor. That does not mean the method will predict labels of new points better. Choosing $k$ by training error alone would reward memorizing familiar examples.
Why zero error is not a victory yet
Overfitting and underfitting
Overfitting occurs when a model responds too strongly to outliers or accidental patterns in the training set and performs worse on new objects. An outlier can have unusual features or an unusual label: for example, a wealthy borrower who still defaulted. A winding boundary may carefully accommodate even that point, but such care with familiar data promises nothing about the next customer.
Underfitting is the opposite problem: the model is not flexible enough to express useful patterns. Generalization is the ability to maintain quality on objects the model did not see during training. We want a boundary that notices meaningful differences without tracing every accident of the sample. Validation helps us choose that balance.
Choosing the number of neighbors
The lecture compares $k = 1$ with $k = 15$: the first gives a boundary that is too jagged, while the second gives a smoother one. On the test-error curve shown there, the minimum occurs at$k = 7$. That value belongs to the example; it is not a universal setting for the method. In your own task, compare candidates on a validation set and reserve the test set for the final evaluation of the chosen candidate.
When everyone can see the test score
Public and private competition sets
In competitions such as Kaggle, a public portion of the test data provides an interim score, while a private portion determines the final result. Repeatedly changing a model based on the public score can favor a candidate that happened to do well on that portion. The private portion helps keep the final score from depending on that search. Even a leaderboard does not remove the need to ask which data guided each decision.
A practical order of work
- Define the object, its features, and the label you want to predict.
- Split labeled data into training and validation sets.
- Train candidates on the training set; select the model and hyperparameters by validation error.
- If a test set is available, evaluate the selected model on it once.
We can now return to the new customer from the opening. Predicting the class is only half the job. The other half is checking that the rule can answer for people it has not met before.
Source linked by the notes: lecture material.

