Skip to main content
←Engineering Blog
Artificial Intelligence2026-08-23

What Actually Happens Inside an Image Segmentation Model?

QV
Quan Van•24 min read

What Actually Happens Inside an Image Segmentation Model?

The Mathematics Behind Image Segmentation

Senior AI Engineer Question

"A segmentation model outputs a mask. But mathematically, how does an image become a mask?"

Most explanations stop at:

Image
 ↓
Segmentation Model
 ↓
Mask

That's useful for beginners.

But it hides almost everything interesting.

A segmentation model is fundamentally a function:

f : R^(H × W × C) → R^(H × W × K)

where:

H = image height
W = image width
C = input channels
K = number of classes

For an RGB image:

C = 3

If we want to segment:

background
person
car

then:

K = 3

The model is therefore not predicting:

one class

for the entire image.

It is producing something closer to:

For every pixel:
    probability distribution over classes

That single idea explains much of segmentation.

Let's derive the entire pipeline.


1. Start With the Image

An RGB image can be represented as a tensor:

X ∈ R^(H × W × 3)

For example:

512 × 512 × 3

means:

512 × 512 pixels
3 channels

A single pixel is:

X[i, j] =
[
    R
    G
    B
]

For example:

X[100, 200] =
[
    0.8
    0.2
    0.1
]

after normalization.

So the segmentation problem starts as pure numerical computation.

There is no concept of:

person
car
tree

inside the raw tensor.

There are only numbers.


2. What Is the Ground Truth?

Suppose we have:

Image

and a human annotator produces:

Mask

For semantic segmentation:

Y ∈ {0, 1, ..., K-1}^(H × W)

For example:

Y[i,j] = 0

means:

background

while:

Y[i,j] = 1

means:

person

and:

Y[i,j] = 2

means:

car

Therefore the ground truth is effectively a matrix:

        pixel
        ↓

Y = [ 0 0 0 0 0
      0 1 1 1 0
      0 1 1 1 0
      0 0 0 0 0 ]

The segmentation task is:

Find f such that:

f(X) ≈ Y

3. But the Model Doesn't Directly Output Classes

This is an important distinction.

Suppose there are:

K = 3

classes.

The model might output:

Z ∈ R^(H × W × 3)

where Z contains logits.

For one pixel:

Z[i,j] =
[
    2.1,
    5.7,
    0.4
]

These aren't probabilities yet.

They're raw scores.


4. Softmax Turns Scores Into Probabilities

For multiclass segmentation:

P(class = k | X, i, j)

can be obtained with Softmax:

P_k =
exp(Z_k)
----------------
Σ_j exp(Z_j)

For:

Z = [2.1, 5.7, 0.4]

we calculate:

P_0 =
e^2.1 / (e^2.1 + e^5.7 + e^0.4)

P_1 =
e^5.7 / (e^2.1 + e^5.7 + e^0.4)

P_2 =
e^0.4 / (e^2.1 + e^5.7 + e^0.4)

The result might look approximately like:

[
    0.027,
    0.962,
    0.011
]

Now the model is saying:

background = 2.7%
person     = 96.2%
car        = 1.1%

for that pixel.


5. Segmentation Is Basically Classification Per Pixel

This is one of the most useful mental models.

Classification:

Image
 ↓
one probability vector

Segmentation:

Image
 ↓
probability vector
 ↓
for every pixel

Therefore:

Classification:
P(class | image)

Segmentation:
P(class | image, pixel)

But there is a huge catch.

Pixels aren't independent.

And that's where the architecture becomes important.


6. Why Can't We Just Classify Every Pixel Independently?

Suppose we did:

pixel 1 → classifier
pixel 2 → classifier
pixel 3 → classifier
...

We would lose spatial context.

Imagine:

red pixel

What is it?

Could be:

apple
car
shirt
wall

The RGB value alone isn't enough.

The surrounding pixels matter.

Therefore the model needs:

local context
+
global context

This is where convolution comes in.


7. Convolution Is a Mathematical Neighborhood Operator

Consider a small image:

X =
[ x11 x12 x13
  x21 x22 x23
  x31 x32 x33 ]

and a kernel:

K =
[ k11 k12 k13
  k21 k22 k23
  k31 k32 k33 ]

The convolution-like operation computes:

Y(i,j)
=
Σ_a Σ_b
K(a,b) X(i+a,j+b)

In plain language:

Take a local neighborhood, multiply each value by a learned weight, and sum the results.

For one location:

Y =
x11*k11 +
x12*k12 +
x13*k13 +
x21*k21 +
x22*k22 +
x23*k23 +
x31*k31 +
x32*k32 +
x33*k33

That is the basic mathematical operation behind CNNs.


8. What Does the Kernel Learn?

Initially:

K

contains numerical parameters.

During training, optimization changes them.

Eventually different kernels can respond strongly to different patterns.

For example, conceptually:

Kernel A → vertical edges
Kernel B → horizontal edges
Kernel C → texture
Kernel D → corners
Kernel E → more complex structures

The network isn't explicitly told:

"learn edges"

The optimization process discovers useful filters.


9. Why Does Convolution Help Segmentation?

Because segmentation depends heavily on spatial structure.

Suppose:

pixel P

is ambiguous.

Its local neighborhood might reveal:

hair
face
shirt
background

The convolution allows the model to calculate features from neighborhoods:

pixel
 ↓
local neighborhood
 ↓
feature

Stack many layers:

local features
 ↓
edges
 ↓
textures
 ↓
parts
 ↓
objects

The receptive field grows.


10. Receptive Field

The receptive field is essentially:

How much of the original image can influence a particular feature.

A shallow layer may see:

3 × 3 pixels

A deeper layer may indirectly depend on:

large regions

Eventually a feature may incorporate information from a significant portion of the image.

Conceptually:

Pixel
 ↓
3×3 neighborhood
 ↓
larger neighborhood
 ↓
object part
 ↓
object
 ↓
global context

This matters because segmentation needs both:

"What is this?"

and:

"Where exactly is it?"

11. The Fundamental Segmentation Conflict

There is a fundamental tension:

High resolution

helps:

precise boundaries

while:

Low resolution

helps:

large receptive fields
semantic understanding
computational efficiency

This is one of the central architectural problems in segmentation.


12. Downsampling

Suppose the original image is:

512 × 512

A CNN may reduce it:

512 × 512
 ↓
256 × 256
 ↓
128 × 128
 ↓
64 × 64
 ↓
32 × 32

Why?

Because lower spatial resolution makes computation cheaper and allows deeper features to represent larger areas.

But something is lost.

A tiny object might occupy:

3 × 3 pixels

at the original resolution.

After repeated downsampling it may become almost nothing.

So:

semantic understanding ↑
spatial precision ↓

This is the segmentation problem.


13. U-Net's Core Idea

U-Net became famous because it addresses this problem with:

Encoder
+
Decoder
+
Skip Connections

Conceptually:

Input
  │
  ▼
Encoder
  │
  ▼
Bottleneck
  │
  ▼
Decoder
  │
  ▼
Mask

But skip connections create another path:

Encoder ───────────────► Decoder

So information from high-resolution layers can be reused.


14. Why Skip Connections Matter Mathematically

Suppose:

F_low

contains:

high spatial resolution

and:

F_high

contains:

high semantic abstraction

The decoder can combine them:

F =
Concat(F_low, Upsample(F_high))

or another fusion operation.

This gives the decoder:

semantic information
+
fine spatial information

That combination is extremely valuable for segmentation.


15. Upsampling

Suppose the decoder has:

64 × 64

but we need:

512 × 512

We need to increase spatial resolution.

Methods include:

nearest-neighbor interpolation
bilinear interpolation
transposed convolution
learned upsampling

16. Bilinear Interpolation Is Pure Mathematics

Suppose we want the value at a new coordinate:

(x, y)

We interpolate using neighboring pixels.

In one dimension:

f(x)
=
(1-t)f(x0)
+
t f(x1)

In two dimensions, bilinear interpolation combines four neighbors.

Conceptually:

Q11 ───── Q21
 │         │
 │    P    │
 │         │
Q12 ───── Q22

Then:

P
=
weighted combination of
Q11, Q12, Q21, Q22

No AI magic.

Just interpolation.


17. Transposed Convolution

Another method is:

ConvTranspose

Conceptually, instead of:

large → small

we learn a transformation that can produce:

small → larger

But it is important not to think of it as a literal mathematical inverse of convolution.

It is better understood as:

A learned upsampling operator.

It contains learnable parameters just like ordinary convolution.


18. Eventually We Reach the Logits

After the encoder and decoder:

Feature Map
 ↓
1×1 convolution
 ↓
K channels

Suppose:

K = 3

Then:

Z ∈ R^(H × W × 3)

For every pixel:

Z[i,j] =
[
    background_score,
    person_score,
    car_score
]

Then:

Softmax

turns them into probabilities.


19. The Final Mask

For semantic segmentation:

ŷ(i,j)
=
argmax_k P_k(i,j)

For every pixel:

choose class with maximum probability

So:

P =
[
  [0.9, 0.1, 0.0],
  [0.1, 0.8, 0.1],
  [0.0, 0.2, 0.8]
]

becomes:

Mask =
[
  0 0 0
  1 1 1
  2 2 2
]

The mask is simply the result of an argmax.


20. Where Does Learning Actually Happen?

Everything we've discussed so far can be deterministic:

convolution
+
activation
+
pooling
+
upsampling
+
softmax
+
argmax

The learned part is the parameters.

For example:

K₁
K₂
K₃
...
W
b

Training tries to find parameters:

θ

such that:

f(X; θ)

matches:

Y

21. Cross-Entropy Loss

For one pixel with true class:

y = 1

and predicted probability:

p = 0.8

cross-entropy is:

L = -log(p)

Therefore:

L = -log(0.8)

which is relatively small.

If:

p = 0.01

then:

L = -log(0.01)

which is very large.

So the loss strongly penalizes confident wrong predictions.


22. Pixel-Wise Cross Entropy

For an entire image:

L =
- Σ_i Σ_j Σ_k
Y[i,j,k] log P[i,j,k]

If the labels are one-hot encoded:

Y[i,j,k] ∈ {0,1}

only the correct class contributes to the sum.

Usually we average:

L =
-1/(HW)
Σ_i Σ_j
log P(correct_class | i,j)

Now the model is being asked:

For every pixel, assign high probability to the correct class.


23. But Cross-Entropy Has a Problem

Suppose:

background = 95%
object = 5%

Then a stupid model could predict:

background everywhere

and still achieve:

95% pixel accuracy

That sounds good.

But the segmentation is useless.

This is called:

Class imbalance.

And it leads to another important mathematical objective.


24. Intersection over Union

IoU is:

IoU =
Intersection
-------------
Union

For predicted mask P and ground truth G:

IoU =
|P ∩ G|
-----------
|P ∪ G|

Suppose:

Intersection = 80
Union = 100

then:

IoU = 0.8

or:

80%

25. Dice Coefficient

Dice is:

Dice =
2|P ∩ G|
-----------
|P| + |G|

Suppose:

|P ∩ G| = 80
|P| = 90
|G| = 100

Then:

Dice =
160 / 190
≈ 0.842

Dice is particularly useful when the foreground occupies a small portion of the image.


26. Soft Dice

The problem with binary masks is that:

argmax

is not nicely differentiable.

During training we want a smooth function.

So we can use probabilities:

p_i

instead of binary predictions.

A common soft Dice form is:

Dice =
2 Σ_i p_i g_i + ε
--------------------------------
Σ_i p_i + Σ_i g_i + ε

where:

ε

prevents division by zero.

Now the loss can be differentiated.

For example:

L_Dice = 1 - Dice

27. Why Dice Is Interesting

Cross-entropy asks:

"Is each pixel classified correctly?"

Dice asks something closer to:

"Does the predicted region overlap the target region?"

These are different objectives.

That's why segmentation models often use combinations such as:

L =
λ₁ L_CE
+
λ₂ L_Dice

Now the model simultaneously cares about:

pixel-level classification
+
region-level overlap

28. IoU and Dice Are Related

For binary sets:

IoU = I / U

and:

Dice = 2I / (P + G)

Since:

U = P + G - I

we can derive:

Dice =
2 IoU
-----------
1 + IoU

and therefore:

IoU =
Dice
--------
2 - Dice

So these metrics are mathematically related.


29. Binary Segmentation Uses Sigmoid

If there are only:

background
foreground

we don't necessarily need Softmax.

The model can output:

z(i,j)

and use:

σ(z)
=
1 / (1 + e^-z)

This produces:

0 < p < 1

representing:

P(foreground)

Then:

p > 0.5

might become:

foreground

otherwise:

background

30. Thresholding Is Another Mathematical Decision

Suppose:

p =
0.1
0.4
0.51
0.9

with threshold:

τ = 0.5

then:

mask =
0
0
1
1

But why exactly:

0.5

?

There is no universal law that says the threshold must be 0.5.

Depending on the application, we might choose:

τ = 0.3

or:

τ = 0.7

to optimize:

precision
recall
IoU
F1

This becomes a decision-theoretic problem.


31. Segmentation Is Not Just One Problem

"Segmentation" actually describes several tasks.

Semantic Segmentation

Every pixel receives a class:

person
person
person
car
car
background

Two people may receive the same class.


Instance Segmentation

Different objects of the same class receive different identities:

Person #1
Person #2
Person #3

Now the problem is:

class
+
instance identity

Panoptic Segmentation

Combines:

semantic segmentation
+
instance segmentation

Now every pixel gets:

class
+
instance ID

This is significantly more complex.


32. Why Instance Segmentation Is Harder

Imagine:

Person A
██████

Person B
██████

Semantic segmentation only needs:

person

for both.

Instance segmentation needs:

Person A → ID 1
Person B → ID 2

So the model needs to reason about object separation.

This introduces another mathematical problem:

How do we distinguish two connected regions belonging to the same semantic class?


33. One Mathematical Approach: Embeddings

Suppose the model maps every pixel to an embedding:

e_i ∈ R^d

Pixels belonging to the same object should have similar embeddings.

Pixels belonging to different objects should be far apart.

We can define:

distance(e_i, e_j)

using Euclidean distance:

d(e_i,e_j)
=
√Σ_k (e_ik - e_jk)^2

Then clustering can separate instances.

Now segmentation becomes partly:

classification
+
metric learning
+
clustering

34. Another Approach: Mask Proposals

Another family of architectures predicts:

candidate object
+
class
+
mask

Conceptually:

Image
 ↓
Feature extraction
 ↓
Object candidates
 ↓
Mask prediction
 ↓
Mask refinement

The mathematical machinery can include:

classification loss
+
bounding-box regression
+
mask loss

So instance segmentation becomes a multi-objective optimization problem.


35. Why Boundaries Are So Difficult

Suppose the ground truth is:

████████
████████
████████

but the prediction is:

███████
███████
███████

Only a few pixels differ.

But those pixels may represent the entire boundary.

Small localization errors can dramatically change:

IoU
Dice
boundary quality

This is why high-resolution information matters.


36. Boundary Accuracy Is Different From Region Accuracy

Consider:

Prediction A:
correct region, slightly wrong boundary

Prediction B:
wrong interior, correct approximate boundary

A simple pixel loss may not capture all the semantic differences humans care about.

Therefore some segmentation systems introduce:

boundary losses
contour losses
distance-transform losses
Hausdorff-related objectives

The objective becomes more geometric.


37. Distance Transform

For a mask, we can calculate the distance from each pixel to the nearest boundary.

Conceptually:

Boundary:
████████

Distance:
0 0 0 0 0

Inside:
1 2 3 2 1

The distance transform provides geometric information.

This can be used to construct losses that penalize boundary errors differently from interior errors.

Now we're no longer thinking purely in terms of:

class probability

but:

geometry

38. Segmentation Is Ultimately a Geometry Problem

This is one of the deeper insights.

Classification asks:

"What is this?"

Segmentation asks:

"What is this AND where exactly does it exist?"

Therefore segmentation involves:

probability
+
spatial structure
+
geometry
+
optimization

That's why the architecture needs both:

semantic abstraction

and:

spatial precision

39. The Full Mathematical Pipeline

We can now summarize a typical semantic segmentation model:

Image
X ∈ R^(H×W×C)
        │
        ▼
Convolution
        │
        ▼
Non-linear transformation
        │
        ▼
Downsampling
        │
        ▼
Deep feature representation
        │
        ▼
Upsampling
        │
        ▼
Skip / Feature Fusion
        │
        ▼
Logits
Z ∈ R^(H×W×K)
        │
        ▼
Softmax
        │
        ▼
P(class | pixel)
        │
        ▼
Argmax
        │
        ▼
Segmentation Mask

During training:

Prediction
     │
     ▼
Loss
     │
     ▼
Gradient
     │
     ▼
Parameter Update
     │
     ▼
Better Prediction

40. Where Does Backpropagation Enter?

Suppose:

L(θ)

is the segmentation loss.

We want:

θ*
=
argmin_θ L(θ)

But directly solving this optimization problem is usually impossible because:

θ

can contain millions or billions of parameters.

So we use gradient-based optimization.

Calculate:

∇θ L

Then update:

θ_{t+1}
=
θ_t
-
η ∇θ L

where:

η

is the learning rate.


41. What Does the Gradient Actually Mean?

The gradient:

∂L / ∂θ

answers:

"If I slightly change this parameter, how does the loss change?"

If:

∂L/∂θ > 0

increasing θ tends to increase the loss locally.

If:

∂L/∂θ < 0

increasing θ tends to decrease the loss locally.

So gradient descent moves parameters toward lower loss.


42. Why Does This Produce a Segmentation Model?

Because the network is a composition:

f(x)
=
f_n(
    f_{n-1}(
        ...
        f_2(
            f_1(x)
        )
    )
)

The chain rule lets us compute:

∂L/∂θ

through the entire computation graph.

This is backpropagation.

Mathematically:

∂L/∂x
=
∂L/∂f_n
·
∂f_n/∂f_{n-1}
·
...
·
∂f_1/∂x

The network learns because calculus tells us how changing parameters changes the objective.


43. There Is No Magic "Segmentation Neuron"

This is important.

There isn't a magical neuron that says:

"This is a person."

Instead:

millions of numerical transformations

collectively produce:

logit(person)

for each pixel.

The model learns a high-dimensional function:

f(X; θ)

that maps:

image

to:

pixel-wise predictions

44. What Is the Model Really Learning?

At a very abstract level:

θ*
=
argmin_θ
E[(X,Y)~D]
[
L(f(X;θ), Y)
]

It is searching for parameters that minimize expected loss over a data distribution.

So the learned model approximates:

P(Y | X)

or a related decision function.

That's the mathematical heart of segmentation learning.


45. But What If We Remove Training?

Now connect this to the previous article.

Suppose we don't train.

We can still build:

edge detection
+
color clustering
+
texture analysis
+
region growing
+
graph cuts
+
watershed
+
thresholding
+
morphological operations

These are classical mathematical segmentation techniques.


46. Classical Segmentation: Thresholding

Suppose grayscale intensity is:

I(x,y)

We choose threshold:

T

and define:

M(x,y)
=
1, if I(x,y) > T
0, otherwise

That's segmentation.

No training.

No neural network.

No GPU.

Just mathematics.


47. Otsu's Method

Instead of manually choosing:

T

Otsu's method chooses a threshold that maximizes between-class variance.

Suppose pixels are divided into:

class 0
class 1

For threshold t, calculate:

σ_B²(t)

the between-class variance.

Then choose:

t*
=
argmax_t σ_B²(t)

Again:

no training

Yet it performs segmentation.


48. Region Growing

Start from a seed:

S

and examine neighboring pixels.

If:

distance(feature(pixel), feature(region))
< τ

then include the pixel.

Repeat.

Eventually:

seed
 ↓
neighbors
 ↓
neighbors
 ↓
region

This is essentially graph traversal under a similarity constraint.


49. Graph Cuts

We can model an image as a graph:

pixel = node
neighbor relationship = edge

Then define an energy:

E(x)
=
E_data(x)
+
λ E_smooth(x)

where:

E_data

measures how well a pixel belongs to a class.

and:

E_smooth

encourages neighboring pixels to have compatible labels.

Then solve:

x*
=
argmin_x E(x)

That's segmentation.

Pure optimization.


50. Watershed

Imagine the image as a topographic surface:

high intensity → mountain
low intensity → valley

Imagine pouring water into the valleys.

Different basins expand until they meet.

The resulting boundaries become segmentation boundaries.

This sounds visual, but underneath it is:

gradient
+
local minima
+
graph propagation

Again:

mathematics

51. Morphological Operations

After segmentation, we may need to clean the mask.

Two fundamental operations are:

erosion
dilation

Given a structuring element B:

erosion:
A ⊖ B

and:

dilation:
A ⊕ B

Then:

opening
=
erosion + dilation

and:

closing
=
dilation + erosion

These can remove noise, fill gaps, and smooth boundaries.


52. This Gives Us an Important Insight

A modern segmentation pipeline might look like:

Neural Model
     ↓
Probability Mask
     ↓
Threshold
     ↓
Morphological Cleanup
     ↓
Connected Components
     ↓
Final Mask

Notice:

Neural network

is only one component.

The rest can be classical mathematics.

This is common in serious computer vision systems.


53. Why Would We Still Use Classical Mathematics?

Because it gives us:

predictability
interpretability
speed
low memory usage
domain-specific control

Suppose an industrial camera detects defects on a metal surface.

The environment might be:

controlled lighting
fixed camera
fixed object
fixed geometry

A carefully designed mathematical pipeline can outperform a huge neural network in:

cost
latency
simplicity

for that specific problem.


54. But Modern Segmentation Wins in Uncontrolled Environments

Imagine:

street scene

with:

different lighting
weather
occlusion
perspective
object shapes
backgrounds
camera angles

Now handcrafted thresholds become fragile.

For example:

person pixel color

can vary enormously.

Learned representations become much more useful.


55. This Is the Fundamental Trade-Off Again

Classical segmentation:

Explicit assumptions
+
Mathematical rules

Modern segmentation:

Learned representation
+
Optimization

Classical methods ask:

"What mathematical rule separates these regions?"

Deep models ask:

"Can I learn a representation in which these regions become separable?"

That's a profound difference.


56. A Senior-Level Mental Model

When you see a segmentation architecture such as:

U-Net
DeepLab
Mask R-CNN
SegFormer
SAM

don't memorize the architecture first.

Ask:

1. What is the input representation?

2. How is spatial information transformed?

3. Where does downsampling happen?

4. How is receptive field increased?

5. How is spatial resolution recovered?

6. How are features fused?

7. What exactly is the output tensor?

8. Which activation converts logits into probabilities?

9. What loss is optimized?

10. How does the loss handle class imbalance?

11. How are boundaries represented?

12. How is the final discrete mask produced?

Once you understand those questions, different architectures become variations on the same mathematical theme.


57. The Entire Problem in One Equation

At a very high level:

ŷ = argmax_y P(y | X; θ)

where:

X

is the image.

The model learns:

θ

such that:

P(y | X; θ)

approximates the correct pixel labeling.

Training solves approximately:

θ*
=
argmin_θ
E[
L(
f(X;θ),
Y
)
]

And inference becomes:

X
 ↓
f(X; θ*)
 ↓
probabilities
 ↓
decision
 ↓
mask

That's the mathematical skeleton behind a huge family of segmentation models.


58. The Most Important Insight

A segmentation model isn't fundamentally:

"an AI that draws masks."

Mathematically, it is:

A parameterized function that transforms a high-dimensional spatial tensor into a dense field of class probabilities, optimized under a geometric and statistical objective.

Once you understand that sentence, the architecture becomes much less mysterious.


Interview Question

"Why does segmentation require an encoder-decoder architecture instead of simply classifying every pixel independently?"

Strong Senior Answer

"Because segmentation requires both semantic context and precise spatial localization. Independent pixel classification loses the spatial relationships between neighboring pixels and cannot effectively distinguish ambiguous pixels based on surrounding context.

An encoder progressively reduces spatial resolution while increasing the semantic receptive field, allowing the network to capture larger contextual structures. However, downsampling loses fine-grained spatial information, which is critical for object boundaries. The decoder therefore reconstructs higher-resolution representations, often using skip connections to combine deep semantic features with earlier high-resolution features.

Mathematically, the network is learning a dense mapping from an input tensor R^(H×W×C) to logits R^(H×W×K). Softmax or sigmoid converts those logits into pixel-wise probabilities, and the final mask is obtained through a decision operation such as argmax or thresholding."


Final Takeaways

1. An image is fundamentally a tensor of numbers.

2. A segmentation model maps that tensor to a dense spatial tensor of logits.

3. Softmax converts multiclass logits into probabilities.

4. Sigmoid is commonly used for binary segmentation.

5. Segmentation can be viewed as classification performed at every pixel.

6. Pixels cannot be treated independently because spatial context matters.

7. Convolution aggregates local neighborhoods.

8. Deep layers increase the effective receptive field.

9. Downsampling increases semantic context but loses spatial precision.

10. Decoder structures recover spatial resolution.

11. Skip connections preserve high-resolution information.

12. Cross-entropy optimizes pixel-level classification.

13. Dice and IoU optimize region overlap more directly.

14. Class imbalance can make raw pixel accuracy misleading.

15. Boundary quality introduces a geometric component to the problem.

16. Instance segmentation adds object identity on top of semantic classification.

17. Embeddings can transform instance separation into a metric-space problem.

18. Classical segmentation can be performed entirely with mathematics.

19. Thresholding, Otsu, watershed, region growing, graph cuts, and morphology do not require neural networks.

20. Modern segmentation models mainly differ in how they construct useful representations and recover spatial information.

21. Training does not "draw the mask"; it learns parameters for a function that produces the mask.

22. Backpropagation computes how changing parameters affects the segmentation loss.

23. The real mathematical problem is:

    θ* = argminθ E[L(f(X;θ),Y)]

24. The final segmentation is essentially a dense decision field:

    pixel → probability distribution → class/mask

25. Understanding this mathematical pipeline is more valuable than memorizing individual model architectures.

One Sentence to Remember

Segmentation is the mathematics of turning spatial evidence into a dense probability field: the model must learn not only what each pixel is, but how pixels relate to one another in space.

EXPLORE THE MANUAL

Want to integrate PomaiDB into your project?

View PomaiDB on GitHub→