An interactive walk through $\sigma(wx+b)$: what $w$ does, what $b$ does, and why this matters for training (Module 3.3) and for building “towers” (Module 3.5).
1The sigmoid neuron
A perceptron outputs a hard 0/1 at the threshold. The sigmoid (logistic) neuron replaces that jump with a smooth S-curve p. 8–10:
$$y = \sigma(z) = \frac{1}{1+e^{-z}}, \qquad z = wx + b \quad\Big(\text{in the slides: } z = w_0 + \textstyle\sum_{i=1}^n w_i x_i,\; b = w_0 = -\theta\Big)$$
The output is always between 0 and 1, so it can be read as a probability (“probability of liking the movie”).
It is smooth, continuous and differentiable everywhere, which is what gradient descent needs. The perceptron's step is not differentiable at the threshold p. 10.
Fixed facts: $\sigma(0)=0.5$, $\sigma(z)\to1$ as $z\to\infty$, $\sigma(z)\to0$ as $z\to-\infty$, and $\sigma'(z)=\sigma(z)(1-\sigma(z))$.
Key idea: the curve $\sigma(z)$ never changes. $w$ and $b$ only change how the input $x$ is mapped to $z$. So they stretch / squeeze / flip ($w$) and slide ($b$) the same S-shape along the x-axis.
2Playground: drag w and b
The grey dashed curve is the reference $\sigma(x)$ ($w=1$, $b=0$). The blue curve is $\sigma(wx+b)$. The orange dot marks the midpoint, where the output equals 0.5.
Midpoint x = −b/w
Slope at midpoint = w/4
Transition width ≈ 4.4/|w|
Output at x=0 = σ(b)
3What the weight w does: steepness and direction
The weight multiplies $x$, so it scales the x-axis. A bigger $|w|$ makes $z$ change faster as $x$ moves, so the curve rises faster.
Value of w
Effect on the curve
$w$ large and positive (e.g. 10, 50)
Very steep. It approaches the step functionp. 56
$w = 1$
The standard sigmoid
$0 < w < 1$
Gentle, stretched-out S. Nearly linear over a wide range
$w = 0$
A flat horizontal line at height $\sigma(b)$. The input is ignored
$w < 0$
Flipped: it goes from 1 down to 0 (mirror image)
Two exact numbers that show this:
$$\left.\frac{d}{dx}\sigma(wx+b)\right|_{\text{midpoint}} = w\,\sigma(0)(1-\sigma(0)) = \frac{w}{4}, \qquad \text{width of the } 10\%\to90\% \text{ rise} = \frac{2\ln 9}{|w|} \approx \frac{4.4}{|w|}$$
Note: if $b \neq 0$, changing $w$ also moves the midpoint, because the midpoint is $-b/w$. With $b=0$ every curve passes through $(0, 0.5)$ and only its steepness changes, as in the plot above.
4What the bias b does: shift left or right
The bias adds a constant to $z$. It doesn't change the steepness. It slides the whole curve along the x-axis.
Increase b → the curve moves LEFT. The neuron “fires” (output > 0.5) for smaller $x$, so it is easier to activate.
Decrease b (more negative) → the curve moves RIGHT. It needs a larger $x$ to fire.
The shift is $\Delta x = -\Delta b / w$. With a steep curve (large $w$), the same change in $b$ moves it less.
This matches the perceptron view p. 5, 8: $b = w_0 = -\theta$. A higher threshold θ means a more negative b, so the curve moves right and the neuron is harder to fire.
Common confusion: “$+b$ shifts right.” It does the opposite. $\sigma(x+3)$ reaches 0.5 at $x=-3$ (left). Rewrite it as $\sigma\big(w(x - x_0)\big)$ with $x_0 = -b/w$ to see the shift directly.
5The midpoint formula: where does the neuron switch?
Click the snapshots from slide 26 to see how $w$ and $b$ move the curve onto the points. Or drag the sliders and try to beat the loss yourself.
f(0.5) → 0.2
f(2.5) → 0.9
Loss
Read the snapshots from top to bottom: w grows (the curve gets steeper, so it can climb from 0.2 to 0.9 between x=0.5 and x=2.5) and b becomes more negative (the curve slides right so the 0.5 crossing falls between the points, at $x_0=2.27/1.78\approx1.28$). Gradient descent p. 41–46 automates exactly these nudges using
Note the extra $x$ in $\nabla_w$: the weight's gradient scales with the input, but the bias's does not, because $\partial z/\partial b = 1$.
7Large w gives a step; two steps give a tower
Module 3.5 p. 56: “If we take the logistic function and set w to a very high value we will recover the step function.” Then $b$ chooses where the step sits ($x_0=-b/w$). Subtract two steps placed at different points and you get a towerp. 56–58:
$$h_{21} = h_{11} - h_{12} = \sigma(w x + b_1) - \sigma(w x + b_2), \qquad \text{tower from } -\tfrac{b_1}{w} \text{ to } -\tfrac{b_2}{w}$$
w controls how sharp the tower walls are. Small $w$ gives a soft bump, large $w$ a crisp rectangle.
b controls where the walls are. The two biases place the left and right edges.
Add many towers, each scaled by an output weight, and you can approximate any function p. 52–54. This is the Universal Approximation Theorem p. 49.
p. 58 asks what activation suits $h_{21}$. It only subtracts, with output weights +1 and −1, so a linear output works.
8Two inputs: $\sigma(w_1x_1 + w_2x_2 + b)$
With two inputs p. 59–64 the sigmoid becomes a 3D surface, and the same rules apply in each direction:
Vary $w_1$ (with $w_2=0$, $b=0$): the surface becomes a step along $x_1$, a cliff parallel to the $x_2$-axis.
Vary $w_2$ (with $w_1=0$): a step along $x_2$.
Vary $b$ (with $w_1=k$ large, $w_2=0$): the cliff slides along $x_1$, to $x_1 = -b/k$.
The boundary where the output is 0.5 is the line $w_1x_1+w_2x_2+b=0$. The vector $(w_1,w_2)$ sets its direction and steepness, and $b$ shifts it.
Subtracting two such steps gives a tower open on two sides p. 61. Combining 4 hidden neurons and passing the result through another sigmoid closes it into a 3D tower p. 62–64 (e.g. $w=200$, $b=\pm100$ on p. 64: huge weights give sharp walls, and the biases place them).
9Cheat sheet
Change
What happens to σ(wx + b)
↑ |w|
Steeper. Approaches a step function
↓ |w| toward 0
Flatter. At $w=0$ it is a constant $\sigma(b)$
w → −w
Curve flips (decreasing)
↑ b
Shifts left by $\Delta b/w$ (fires earlier)
↓ b
Shifts right (fires later, i.e. a higher threshold)
Scale w and b together by k
Same midpoint $-b/w$, but k× steeper
Midpoint (output 0.5)
$x_0=-b/w$
Slope at the midpoint
$w/4$, the maximum slope
Output range
Always $(0,1)$. $w$ and $b$ never change the range
Exam traps: (1) $+b$ shifts left, not right. (2) Changing $w$ also moves the midpoint when $b\neq0$. (3) A negative $w$ flips the curve. It does not shift it. (4) Neither $w$ nor $b$ changes the output range (0, 1).
10Practice questions (click to reveal)
Where does $\sigma(4x - 8)$ cross 0.5, and what is its slope there?
$x_0 = 8/4 = 2$. The slope is $4/4 = 1$.
$\sigma(x)$ and $\sigma(x+5)$: which one is further left?
$\sigma(x+5)$. Its midpoint is at $x=-5$.
You want a neuron that switches sharply from 1 to 0 at $x=3$. Give one choice of w, b.
Use a negative, large $w$: e.g. $w=-20$, $b=60$, since $-b/w = 60/20 = 3$.
What is $\sigma(0\cdot x + b)$ for $b=0$? For $b$ very large?
Constant 0.5 for $b=0$. Constant ≈ 1 for very large $b$. In both cases the input has no effect.
σ(w x + b) with w = 2, b = −4 vs. w = 20, b = −40. What is the same and what is different?
Same midpoint $x_0=2$. The second is 10× steeper (slope 5 vs. 0.5), so it is almost a step.
In the p. 26 training snapshots, why does b become more negative as training proceeds?
The midpoint $-b/w$ must sit between $x=0.5$ and $x=2.5$. As $w$ grows to make the curve steeper, $b$ must become more negative to keep the midpoint around $x\approx1.3$.
Make a tower of height 1 between $x=1$ and $x=4$ using two sigmoids with $w=50$.
$\sigma(50x - 50) - \sigma(50x - 200)$, i.e. $b_1=-50$ and $b_2=-200$.