B4B33RPZ · Pattern Recognition and Machine Learning

RPZ: Introduction

Lab 1: course logistics, expectations & a probability recap
Klára JanouškováFEE CTU, Prague

About me

klara.janouskova@fel.cvut.cz

4th year computer vision PhD student
Visual Recognition Group, prof. Jiří Matas

  • Bc and Ing from CTU (Open Informatics)

Experience

  • UPV Valencia (Spain)
  • UAB Barcelona (Spain)
  • IBM Research Zurich
  • Technion (Haifa, Israel)
Latest work: at the intersection of computer vision, reinforcement learning and robotics.
multimodal fine-grained recognition
dataset quality: ImageNet
image region tokens masked pooling ↔ “a red mug”
region embeddings from VLMs & MLLM annotations (SEAR-VLM)
reinforcement learning for object tracking
Student project success story :)

From a student project to top-tier venues

4top-tier publications
~2years

Nikita and Illia both started with a student project in our group.

  • Flaws of ImageNet, Computer Vision's Favorite Dataset
    ICLR 2025 · Singapore (Illia)
  • Image Recognition with Vision and Language Embeddings of VLMs
    BMVC 2025 · UK (Nikita)
  • Multimodal Large Language Models as Image Classifiers
    CVPR 2026 Findings · Denver (Illia)
  • Doomed to Re-Annotate, Forever: The ImageNet Story
    NeurIPS 2026 · Sydney (Nikita)

What you will learn

Part 1

Decision making

  • Bayesian decision theory
  • Non-Bayesian tasks
Part 2

Parameter estimation

  • Parameter estimation of prob. models
  • Nearest neighbour, non-parametric density estimation
Part 3

Classifiers, learning

  • Logistic regression
  • Perceptron
  • SVM
  • AdaBoost
  • Neural networks
  • K-means
  • Expectation Maximization
  • Feature selection & extraction (PCA, …)
  • Decision trees

Lab: homework

You will implement:

02Bayesian decision task
03Non-Bayesian tasks: the minimax task
04Non-parametric estimates: Parzen windows
05MLE, MAP and Bayes parameter estimation
06Logistic regression
07Linear classifier: Perceptron
08Support Vector Machines
09AdaBoost
10K-means clustering
11Convolutional neural networks
Part I

RPZ expectations, requirements, grading, general advice

RPZ: Expectations

Study plan: 156 hours (6 credits)

  • 1 hour/week working through lecture materials
  • 11 lab assignments ≈ 6–7 hours/week
Please benchmark your time on the lab assignments!

Skills

  • Programming in Python (numpy, Jupyter notebooks)
  • Mathematics
    • probability and statistics
    • derivatives, integrals
    • linear algebra, optimization

Requirements, grading

(written exam + tests) is about working with equations and solving exercises.

Semester work (assignments + tests) = 50 %, exam = 50 %.
On top: bonus lab tasks (up to +0.8 p each).

Labs: points

14points for the 11 assignments
+
36points for 3 tests (12 each)
=
50lab points = 50 % of the final grade
Points per assignment
1.3 p
submitted ≤ 2 weeks after release
1 p
later than 2 weeks

First lab (today): 1 point. 10 × 1.3 + 1 = 14.

Start early: being late on all 10 costs 10 × 0.3 = 3 points, a quarter of a test.
Bonuses
  • some assignments have a bonus task: up to +0.8 p
  • show it to your TA at the lab; evaluated individually
  • any time in the semester, at the latest the day before the exam
  • can take you above 50 or make up for lost test points

Labs: requirements (zápočet)

All 11 assignments must pass the BRUTE evaluation by 10 January 2027. After the deadline, no more submissions.

Tests

  • short written tests at the start of the lab on 16.10., 13.11. and 11.12.
  • taking them is required for the zápočet
  • questions refer to the preceding lectures
  • missed test: replacement only in exceptional, well-justified cases, agreed individually

Communication

  • Asking for help: use the forum whenever possible (fastest, all TAs watch it)
  • email your TA for personal matters (can't attend a test, consultations)

Exam

  • zápočet is required to sign up
  • minimum score in the written part (typically 5–10 of 40)
  • the oral exam is mandatory

The semester at a glance

September – November
25.9.
Intro, numpy
2.10.
Bayes decision
9.10.
Minimax
16.10.
Parzen
TEST
23.10.
MLE / MAP
30.10.
Logistic reg.
6.11.
Exercises
13.11.
Perceptron
TEST
November – January
20.11.
Dean's day
27.11.
SVM
4.12.
AdaBoost
11.12.
k-means
TEST
18.12.
CNN
8.1.
Wrap-up
10.1.
Deadline
all 11 in BRUTE
Each assignment is released at its lab: 2 weeks for full points (1.3 p).
Tests at the start of the lab: questions refer to the preceding lectures.

During the labs

The labs are not meant for coding the homework assignments!

We will:

  • practice exercises
    • exercise book on the web page
    • similar to what you can expect in tests/exam
  • deepen understanding of the lecture material
  • introduce the assignments
    • but you are expected to do the coding at home
    • more about discussing and practicing the relevant concepts than going through the assignment itself (depends on the TA)

AI tools, cooperation

Cooperation

Encouraged, but it is NOT ok to:

  • show your implementation or derivation to others
  • use an existing implementation / derivation (except teaching materials and Python/PyTorch tutorials)
  • publish or otherwise share your solutions
Submitting plagiarised code = no zápočet.
AI tools

Suggestion: don't use AI tools (Copilot, ChatGPT, …) to code the assignments

  • not prohibited: use them smartly to learn more, e.g. they may suggest a better implementation than yours
  • avoid “vibe-coding” the assignments
  • you are adults: if you don't learn anything from the assignment, that's your problem

Hands up if you…

No wrong answers: this helps me pitch the labs at the right level.

Coding & ML
●●●●●have used Python✋ –
●●●●●have used numpy✋ –
●●●●●have trained a neural network✋ –
●●●●●know how a transformer works✋ –
●●●●●have heard of a Gated DeltaNet✋ –
Maths
like maths✋ –
hate maths✋ –
AI tools
have used an LLM coding agent✋ –
have a paid AI subscription✋ –

A bit more about you

Why are you taking RPZ?
it is required✋ –
I am curious about ML✋ –
I want to do ML research✋ –
Study programme
Open Informatics✋ –
Cybernetics and Robotics✋ –
something else✋ –
Warm-up

I flip a fair coin 3 times. What is the chance of at least one head?

Complement rule: P(at least one H)=1−P(no H)P(\text{at least one H}) = 1 - P(\text{no H})
=1−(12)3=1−18=78= 1 - \left(\tfrac{1}{2}\right)^3 = 1 - \tfrac{1}{8} = \tfrac{7}{8}
HHHHHTHTHHTTTHHTHTTTHTTT
8 equally likely outcomes; only TTT has no head.
Part II

Assignment 1

Open the assignment page →

Goals

  1. Prepare the coding environment
  2. Verify you have access to the student forum, git templates and BRUTE
  3. If you have not used them before:
    • get familiar with basic git and GitLab usage
    • learn the basics of numpy: do not underestimate this!
  4. Download the git template for the first lab
  5. Verify you can upload to BRUTE
Part III

Recap

Probability: the three axioms

A1 · Non-negativity
P(A)≥0P(A) \ge 0
ΩA
A2 · Normalization
P(Ω)=1P(\Omega) = 1
Ωwhole spacesomething always happens
A3 · Additivity
P(A∪B)=P(A)+P(B)P(A \cup B) = P(A) + P(B)  if  A∩B=∅A \cap B = \emptyset
ΩABno overlap: A ∩ B = ∅
Consequence · overlapping events
P(A∪B)=P(A)+P(B)−P(A∩B)P(A \cup B) = P(A) + P(B) - P(A \cap B)
ΩABA ∩ B is counted twice

Think of probability as area inside Ω. Proper definition: en.wikipedia.org/wiki/Probability_axioms

🐷 🐑 🐺

Pigs and sheep (🐷, 🐑) live peacefully together, divided into two groups ([ ], [ ]).

One day, a wolf comes to the animals.

The wolf is hungry. He will randomly pick an animal and have it for lunch.

Notation: one random pick, many questions
P(🐷): the probability that lunch is a pig
P([ ]): the probability that lunch comes from the blue (solid) pen
P(🐷, [ ]): lunch is a pig and comes from the blue pen
P(🐷 | [ ]): lunch is a pig, given we know it came from the blue pen. “|” reads “given”.
sheeppigsheepsheeppigsheep
pigpigsheeppig
wolf

P(A=a∣B=b)P(A{=}a \mid B{=}b) vs. P(B=b∣A=a)P(B{=}b \mid A{=}a)

P(🐷) =
5/10 = ½
P(🐑) =
5/10 = ½
P([ ]) =
4/10 = ⅖
P([ ]) =
6/10 = ⅗
P(🐷 | [ ]) =
¾
P(🐑 | [ ]) =
4/6 = ⅔
P([ ] | 🐷) =
⅗
P([ ] | 🐑) =
⅘
Definition: P(A∣B)=P(A,B)P(B)P(A \mid B) = \dfrac{P(A, B)}{P(B)}, the share of BB that is also AA.

The wolf is hungry. He will randomly pick an animal and have it for lunch.

sheeppigsheepsheeppigsheep
pigpigsheeppig

Click a box (or press → ) to reveal the answer.

Rules of probability

Think of probability as area inside Ω.

Sum rule

P(A)=P(A,B)+P(A,Bc)P(A) = P(A, B) + P(A, B^c)

Cut A along B: the two pieces add up.

Ω A BBᶜ A, BA, Bᶜ

Total probability (marginalization)

P(A)=∑iP(A,Bi)=∑iP(A∣Bi) P(Bi)P(A) = \sum_i P(A, B_i) = \sum_i P(A \mid B_i)\,P(B_i)

BiB_i: disjoint pieces that cover Ω (exactly one occurs).

Ω AB₁B₂B₃B₄sum of 4 pieces

Product rule

P(A,B)=P(A∣B) P(B)=P(B∣A) P(A)P(A, B) = P(A \mid B)\,P(B) = P(B \mid A)\,P(A)

First land in B, then in A within B: multiply the two chances. E.g. A = 🐷, B = blue pen.

🐷🐷🐷🐑🐷🐷🐑🐑🐑🐑4/10🐷🐷🐷🐑3/4🐷🐷🐷all 10 animalsin B🐷 in BP(🐷, B) = 4/10 · 3/4 = 3/10

More 🐷🐑🐷🐑

P(🐷) = ½
P(🐑) = ½
P([ ]) = ⅖
P([ ]) = ⅗
P(🐷 | [ ]) = ¾
P(🐑 | [ ]) = ⅔
P([ ] | 🐷) = ⅗
P([ ] | 🐑) = ⅘
P(🐷, [ ]) =
P(🐷 | [ ]) P([ ])
= ¾ · ⅖ = 0.3
= P([ ] | 🐷) P(🐷)
= ⅗ · ½ = 0.3
P(🐷, [ ]) =
P(🐷 | [ ]) P([ ])
= ⅓ · ⅗ = 0.2
P(🐷) =
P(🐷, [ ]) + P(🐷, [ ])
= 0.3 + 0.2 = 0.5
sheeppigsheepsheeppigsheep
pigpigsheeppig
wolf

The question Bayes answers

🐺 The wolf's tracks lead to the blue (solid) pen.

Question: was lunch a 🐷?
We want P(🐷∣blue)P(🐷 \mid \text{blue}).
sheeppigsheepsheeppigsheep
pigpigsheeppig

Here we can just count: 3 of the 4. Usually we cannot see everything.

What we typically have
A test is checked on sick patients: we know P(+∣sick)P(+ \mid \text{sick}). We want P(sick∣+)P(\text{sick} \mid +).
In RPZ: we learn how each class looks, P(x∣k)P(x \mid k). We want P(k∣x)P(k \mid x).
Naming the pieces
H, hypothesis: unknown, what we want (sick; class kk; 🐷)
E, evidence: what we observed (test +; features xx; tracks)

Known: P(E∣H)P(E \mid H). Wanted: P(H∣E)P(H \mid E).
How do we flip it?

Bayes' rule: you already know it

1. Product rule, written both ways:

P(H,E)=P(H∣E) P(E)=P(E∣H) P(H)P(H, E) = P(H \mid E)\,P(E) = P(E \mid H)\,P(H)

2. Divide by P(E)P(E):

P(H∣E)=P(E∣H) P(H)P(E)P(H \mid E) = \frac{P(E \mid H)\,P(H)}{P(E)}

Nothing new: Bayes flips the condition, from P(E∣H)P(E \mid H) to P(H∣E)P(H \mid E).

Check on the pens, where we can count: HH = 🐷, EE = blue pen.

P(🐷∣E)=P(E∣🐷) P(🐷)P(E)=35⋅1225=34P(🐷 \mid E) = \frac{P(E \mid 🐷)\,P(🐷)}{P(E)} = \frac{\frac{3}{5} \cdot \frac{1}{2}}{\frac{2}{5}} = \frac{3}{4}

✓ Same as counting: 3 of the 4 animals in the blue pen are pigs.

Updating a belief: prior × likelihood, then rescale

The wolf's tracks lead to the blue pen (evidence EE). Was lunch a 🐷 or a 🐑?
P(H∣E)⏟posterior  ∝  P(E∣H)⏟likelihood  P(H)⏟prior\underbrace{P(H \mid E)}_{\text{posterior}} \;\propto\; \underbrace{P(E \mid H)}_{\text{likelihood}}\;\underbrace{P(H)}_{\text{prior}}

Prior, before the evidence: P(🐷)=P(🐑)=0.5P(🐷) = P(🐑) = 0.5

× Likelihood, how well each explains EE: P(E∣🐷)=0.6P(E \mid 🐷) = 0.6, P(E∣🐑)=0.2P(E \mid 🐑) = 0.2. The bars no longer sum to 1: the likelihood is not a distribution over HH.

÷ P(E), rescale so they sum to 1: P(E)=0.30+0.10=0.40P(E) = 0.30 + 0.10 = 0.40 (total probability).

Bayes' theorem, Bayesian statistics

posteriorP(H∣E)P(H \mid E)
=
priorP(H)P(H)
likelihoodP(E∣H)P(E \mid H)
marginalP(E)P(E)
HH: hypothesis
EE: evidence

Posterior ∝ Prior × Likelihood; P(E)P(E) only rescales so the posterior sums to 1.

Later in RPZ
P(θ∣D)=P(θ) P(D∣θ)P(D)P(\theta \mid D) = P(\theta)\,\frac{P(D \mid \theta)}{P(D)}

θ\theta: model, parameter(s)  ·  DD: data

Often used
P(θ∣D)∝P(D∣θ) P(θ)P(\theta \mid D) \propto P(D \mid \theta)\,P(\theta)

Bayes' theorem: medical example

1 · Formula first
P(sick∣+)=P(+∣sick) P(sick)P(+∣sick) P(sick)+P(+∣healthy) P(healthy)P(\text{sick} \mid +) = \frac{P(+ \mid \text{sick})\,P(\text{sick})}{P(+ \mid \text{sick})\,P(\text{sick}) + P(+ \mid \text{healthy})\,P(\text{healthy})}

Denominator P(+)P(+) = total probability over sick and healthy.

2 · Then the numbers

Prevalence 1 %, sensitivity 90 %, false-positive rate 9 %.

P(sick∣+)=0.9⋅0.010.9⋅0.01+0.09⋅0.99=0.0090.098≈9 %P(\text{sick} \mid +) = \frac{0.9 \cdot 0.01}{0.9 \cdot 0.01 + 0.09 \cdot 0.99} = \frac{0.009}{0.098} \approx 9\,\%
Same thing, counting 1,000 people
10 are sick → 9 test positive
990 are healthy → 89 test positive
P(sick | +) = 9 / (9 + 89) ≈ 9 %

A positive result from a "90 % accurate" test means only about a 1 in 11 chance of being sick: the prior (1 %) matters.

More: Tom Rocks Maths · UPenn practice problems (PDF) · interactive version in the optional section.

Exercise: the lazy short-sighted student

A student at Karlovo náměstí can't read tram numbers, only the tram type. Trams 3, 6, 14, 22, 24 stop there; only 14 and 24 go to Albertov. The joint probability p(x,k)p(x,k) of tram type x∈{old,new}x \in \{\text{old}, \text{new}\} and line kk is:

36142224p(x)p(x)
old0.050.150.100.250.050.60
new0.200.000.050.000.150.40
p(k)p(k)0.250.150.150.250.20
Row sumsp(x)=∑kp(x,k)p(x) = \sum_k p(x,k)
p(old)=0.05+0.15+0.10+0.25+0.05=0.60p(\text{old}) = 0.05 + 0.15 + 0.10 + 0.25 + 0.05 = 0.60
Column sumsp(k)=∑xp(x,k)p(k) = \sum_x p(x,k)
p(3)=0.05+0.20=0.25p(3) = 0.05 + 0.20 = 0.25
Conditionalp(k∣x)=p(x,k) / p(x)p(k \mid x) = p(x,k)\,/\,p(x)
p(14∣old)=0.10/0.60≈0.17p(14 \mid \text{old}) = 0.10 / 0.60 \approx 0.17
Albertovp(A∣x)=p(14∣x)+p(24∣x)p(\text{A} \mid x) = p(14 \mid x) + p(24 \mid x)
p(A∣old)=(0.10+0.05)/0.60=0.25p(\text{A} \mid \text{old}) = (0.10 + 0.05) / 0.60 = 0.25
p(A∣new)=(0.05+0.15)/0.40=0.5p(\text{A} \mid \text{new}) = (0.05 + 0.15) / 0.40 = 0.5

From lecture 1 (J. Matas): Introduction. Bayesian Decision Theory. The lecture continues with losses and the optimal strategy.

Bayes, visually: who is Pepa?

Pepa is shy, tidy and loves order and detail.

Is he more likely a ● librarian or a ● farmer?

P(lib∣E)=P(E∣lib) P(lib)P(E)P(\text{lib} \mid E) = \frac{P(E \mid \text{lib})\,P(\text{lib})}{P(E)}

Prior. For every librarian there are about 20 farmers: say 10 vs. 200.

Evidence. 70 % of librarians fit the description, but only 10 % of farmers do.

Condition on E. Keep only people who fit: 7 librarians and 20 farmers.

Posterior: P(lib∣E)=77+20≈26 %P(\text{lib} \mid E) = \frac{7}{7+20} \approx 26\,\%

Inspired by 3Blue1Brown, Bayes theorem, the geometry of changing beliefs; the example comes from Kahneman & Tversky.

Bayes' theorem as areas

P(H∣E)=P(H) P(E∣H)P(H) P(E∣H)+P(¬H) P(E∣¬H)P(H \mid E) = \frac{P(H)\,P(E \mid H)}{P(H)\,P(E \mid H) + P(\neg H)\,P(E \mid \neg H)}

The square holds all possibilities (total area = 1).

Split by the prior: the thin strip is P(H)=121P(H) = \tfrac{1}{21}, the rest is P(¬H)=2021P(\neg H) = \tfrac{20}{21}.

Shade what fits the evidence: P(E∣H)=0.7P(E \mid H) = 0.7 of the strip, P(E∣¬H)=0.1P(E \mid \neg H) = 0.1 of the rest.

Seeing E restricts us to the shaded area. The posterior is the fraction of it inside HH: numerator = shaded strip, denominator = all shaded.

Evidence doesn't set your belief on its own; it updates the prior.

Why derivatives? We minimise a loss

Training a model = finding parameters θ\theta that make a loss L(θ)L(\theta) (how wrong the model is) as small as possible.

The slope tells us which way is downhill.

At the minimum the curve is flat: slope = 0.

So we need slopes, i.e. derivatives.
θloss L(θ): how wrong the model isslope < 0:go right →slope > 0:← go leftslope = 0: best θ

Derivatives: how fast does f change?

Recall: minimum of the loss = slope 0. How do we measure a slope?

Line: slope =riserun=ΔfΔx= \dfrac{\text{rise}}{\text{run}} = \dfrac{\Delta f}{\Delta x}, the same everywhere (here ½).

run 2rise 1

Curve: the slope changes. Take a second point x0+hx_0 + h: secant slope =f(x0+h)−f(x0)h= \dfrac{f(x_0+h) - f(x_0)}{h}.

Shrink hh: the secant turns towards the direction of the curve at x0x_0.

Derivative = slope of the tangent: f′(x0)=lim⁡h→0f(x0+h)−f(x0)hf'(x_0) = \lim_{h \to 0} \frac{f(x_0+h) - f(x_0)}{h}
f(x) = x² h = 1Δf=3 slope 3 secant slopes 3 → 2.5 → 2.2 tangent: slope 2 x₀ = 1

Worked example: f(x)=x2f(x) = x^2

f′(x0)=lim⁡h→0f(x0+h)−f(x0)hf'(x_0) = \lim_{h \to 0} \frac{f(x_0+h) - f(x_0)}{h}
Plug in
(x0+h)2−x02h\dfrac{(x_0+h)^2 - x_0^2}{h}
Expand
=2x0h+h2h= \dfrac{2x_0 h + h^2}{h}
Simplify
=2x0+h= 2x_0 + h
h→0h \to 0
f′(x0)=2x0f'(x_0) = 2x_0
Check with numbers at x0=1x_0 = 1: secant slope =2+h= 2 + h
hhsecant slope
13
0.52.5
0.12.1
0.012.01
→0\to 02 = f′(1)

Next: try it yourself on any function →

Derivatives: try it yourself

f′(x0)=lim⁡h→0f(x0+h)−f(x0)hf'(x_0) = \lim_{h \to 0} \frac{f(x_0+h) - f(x_0)}{h}
secant slope
tangent slope f′(x0)f'(x_0)

Derivatives: what you'll need in RPZ

Rules
(xn)′=nxn−1(x^n)' = n x^{n-1}(ex)′=ex(e^x)' = e^x
(ax)′=axln⁡a(a^x)' = a^x \ln a (for a=ea = e: ln⁡e=1\ln e = 1, so (ex)′=ex(e^x)' = e^x)
(ln⁡x)′=1x(\ln x)' = \frac{1}{x}(af+bg)′=af′+bg′(a f + b g)' = a f' + b g'
(fg)′=f′g+fg′(f g)' = f' g + f g'(fg)′=f′g−fg′g2\left(\frac{f}{g}\right)' = \frac{f' g - f g'}{g^2}
Chain rule: (f(g(x)))′=f′(g(x)) g′(x)\big(f(g(x))\big)' = f'(g(x))\, g'(x)
Gradient: ∇f(x)=(∂f∂x1,…,∂f∂xn) ⁣⊤\nabla f(\mathbf{x}) = \left(\frac{\partial f}{\partial x_1}, \dots, \frac{\partial f}{\partial x_n}\right)^{\!\top}
Best fit: set the slope to 0 (MLE, MAP)

Maxima and minima have zero slope. Solve

∂∂θlog⁡p(D∣θ)=0\frac{\partial}{\partial \theta} \log p(D \mid \theta) = 0

p(D∣θ)p(D \mid \theta): how well parameters θ\theta explain the data DD (the likelihood).

Can't solve it? Walk downhill (logistic regression, neural nets)
θ←θ−η ∇θL(θ)\theta \leftarrow \theta - \eta\, \nabla_\theta L(\theta)

L(θ)L(\theta): the error we want small. η\eta: step size (learning rate). The chain rule applied layer by layer = backpropagation.

Quick check

Is P(A∣B)=P(B∣A)P(A \mid B) = P(B \mid A)?

Bayes: P(A∣B)=P(B∣A) P(A)P(B)P(A \mid B) = P(B \mid A)\,\frac{P(A)}{P(B)}. They are equal exactly when P(A)=P(B)P(A) = P(B) (both non-zero).

A test detects a disease in 99 % of sick people. You test positive. How likely are you sick?

We know P(+∣sick)P(+ \mid \text{sick}), not P(sick∣+)P(\text{sick} \mid +). We also need the prior P(sick)P(\text{sick}) and the false-positive rate.

Click an answer.

Quick check

A DNA trace matches 1 in a million people. The suspect matches. Is the chance he is innocent 1 in a million?

1 in a million is P(match∣innocent)P(\text{match} \mid \text{innocent}). In a city of 2 million, about 2 innocent people match too. Confusing the two is the prosecutor's fallacy.

Can a probability density p(x)p(x) be larger than 1?

Only the total area must be 1. A Gaussian with σ=0.1\sigma = 0.1 peaks at about 4.

Optional

Extra practice & interactive recap

Use if there is time left, or share the file for self-study.

From probabilities to decisions

The student sees the tram type xx and decides: run or don't run. Losses W(k,d)W(k, d) in CZK:

rundon't run
tram goes to Albertov0
tram does not go there

Expected loss (risk) of decision dd after seeing xx:

R(x,d)=∑kp(k∣x) W(k,d)R(x, d) = \sum_k p(k \mid x)\, W(k, d)

From the table: p(Albertov∣old)=0.25p(\text{Albertov} \mid \text{old}) = 0.25, p(Albertov∣new)=0.5p(\text{Albertov} \mid \text{new}) = 0.5.

Medical test: 1,000 people

P(sick∣+)=sens⋅prevsens⋅prev+(1−spec)(1−prev)P(\text{sick} \mid +) = \frac{\text{sens} \cdot \text{prev}}{\text{sens} \cdot \text{prev} + (1 - \text{spec})(1 - \text{prev})}
sick, test + sick, test − healthy, test + healthy, test −

The Gaussian (normal) distribution

p(x)=1σ2π e−(x−μ)22σ2p(x) = \frac{1}{\sigma\sqrt{2\pi}}\, e^{-\frac{(x-\mu)^2}{2\sigma^2}}

2D Gaussian: what the covariance does

Σ=(σ12ρ σ1σ2ρ σ1σ2σ22)\Sigma = \begin{pmatrix} \sigma_1^2 & \rho\,\sigma_1\sigma_2 \\ \rho\,\sigma_1\sigma_2 & \sigma_2^2 \end{pmatrix}

Ellipses: 1σ, 2σ, 3σ contours. Arrows: eigenvectors of Σ (principal axes; PCA later in the course).

Integrals: probability is area

P(a≤X≤b)=∫abp(x) dx≈∑i=1np(xi) ΔxP(a \le X \le b) = \int_a^b p(x)\,dx \approx \sum_{i=1}^{n} p(x_i)\,\Delta x

Why we always take the log

Likelihood of nn independent samples is a product of small numbers:

L=∏i=1np(xi)⇒log⁡L=∑i=1nlog⁡p(xi)L = \prod_{i=1}^{n} p(x_i) \quad\Rightarrow\quad \log L = \sum_{i=1}^{n} \log p(x_i)

Computed live in float64, the same as numpy's default.

Product
Sum of logs

Extra

3blue1brown.com/topics/probability

Check the videos related to Bayes!