Linear classifiers

Margin

A desk for the hinge, the regularizer, and softmax. Copper is a penalty that is still on. Navy is a constraint that has gone quiet.

01 Multiclass SVM

The correct score has to clear a margin.

Each wrong class contributes max(0, sⱼ − sᵧ + Δ). Once the correct class leads by Δ, that term goes quiet. The loss does not care how much further you win.

Lᵢ = Σⱼ≠yᵢ max(0, sⱼ − sᵧᵢ + Δ)

cat loss
2.90
  • max(0, 5.10 − 3.20 + 1.00) = 2.90
  • max(0, -1.70 − 3.20 + 1.00) = 0.00

The navy dot is this image. Left of the kink, some rival still sits inside the margin of 6.10.

Mean over 3 images
5.27
Lecture check
5.27

With the lecture scores, Δ = 1 and a sum, the three losses are 2.9, 0 and 12.9. Their mean is 5.27. You are on that point.

The car column is the lecture’s quiet example: 1.3 and 2.0 both lose to 4.9 by more than 1, so that image adds nothing. The frog does not — both rivals clear its score.

The lecture’s questions, answered on these numbers

What if the car scores move a little?

If the correct score still leads by at least Δ, every hinge on that image stays 0. The loss is flat inside the margin — the gradient with respect to those scores is zero.

What are the min and max?

Minimum is 0, when the correct class leads every other by Δ. There is no finite maximum: a wrong score can grow without bound and the hinge grows with it.

If the scores were random and roughly equal?

Each of the C−1 terms is about max(0, Δ) = Δ. With Δ = 1 that is about C−1. Here that is about 2; on CIFAR-10 it is about 9. A sane untrained loss.

What if the sum included the correct class?

The j = yᵢ term is max(0, Δ) every time. You add a constant. The W that minimizes the loss does not change. Switch on “Include correct” and watch every column rise by Δ.

What if we averaged the hinges instead of summing?

You divide by C−1. Rankings of models stay the same. The tradeoff against λR(W) does not, so λ would need to be retuned. The argmin over W is unchanged only if λ is scaled to match.

What about a squared hinge?

Still zero inside the margin. A violation of size z now costs z², so large mistakes dominate and small ones barely count. That is a different preference over errors, not a fix for the flat region.

If some W has L = 0, is it unique?

No. f(x, W) = Wx, so 2W doubles every score and every gap. A satisfied margin stays satisfied, and L(2W) = 0 too. Data loss alone will not choose between them. That is why the next station adds R(W).

The formulas and the cat, car, and frog numbers follow Justin Johnson’s linear classifiers lecture, 11 September 2019.