CS6910 · Week 2 (Lecture 3) · Sigmoid Neuron

How the weight and bias reshape the sigmoid

An interactive walk through $\sigma(wx+b)$: what $w$ does, what $b$ does, and why this matters for training (Module 3.3) and for building “towers” (Module 3.5).

1The sigmoid neuron

A perceptron outputs a hard 0/1 at the threshold. The sigmoid (logistic) neuron replaces that jump with a smooth S-curve p. 8–10:

$$y = \sigma(z) = \frac{1}{1+e^{-z}}, \qquad z = wx + b \quad\Big(\text{in the slides: } z = w_0 + \textstyle\sum_{i=1}^n w_i x_i,\; b = w_0 = -\theta\Big)$$
Key idea: the curve $\sigma(z)$ never changes. $w$ and $b$ only change how the input $x$ is mapped to $z$. So they stretch / squeeze / flip ($w$) and slide ($b$) the same S-shape along the x-axis.

2Playground: drag w and b

The grey dashed curve is the reference $\sigma(x)$ ($w=1$, $b=0$). The blue curve is $\sigma(wx+b)$. The orange dot marks the midpoint, where the output equals 0.5.

1.0
0.0
Midpoint x = −b/w
Slope at midpoint = w/4
Transition width ≈ 4.4/|w|
Output at x=0 = σ(b)

3What the weight w does: steepness and direction

The weight multiplies $x$, so it scales the x-axis. A bigger $|w|$ makes $z$ change faster as $x$ moves, so the curve rises faster.

Value of wEffect on the curve
$w$ large and positive (e.g. 10, 50)Very steep. It approaches the step function p. 56
$w = 1$The standard sigmoid
$0 < w < 1$Gentle, stretched-out S. Nearly linear over a wide range
$w = 0$A flat horizontal line at height $\sigma(b)$. The input is ignored
$w < 0$Flipped: it goes from 1 down to 0 (mirror image)

Two exact numbers that show this:

$$\left.\frac{d}{dx}\sigma(wx+b)\right|_{\text{midpoint}} = w\,\sigma(0)(1-\sigma(0)) = \frac{w}{4}, \qquad \text{width of the } 10\%\to90\% \text{ rise} = \frac{2\ln 9}{|w|} \approx \frac{4.4}{|w|}$$

Note: if $b \neq 0$, changing $w$ also moves the midpoint, because the midpoint is $-b/w$. With $b=0$ every curve passes through $(0, 0.5)$ and only its steepness changes, as in the plot above.

4What the bias b does: shift left or right

The bias adds a constant to $z$. It doesn't change the steepness. It slides the whole curve along the x-axis.

This matches the perceptron view p. 5, 8: $b = w_0 = -\theta$. A higher threshold θ means a more negative b, so the curve moves right and the neuron is harder to fire.

Common confusion: “$+b$ shifts right.” It does the opposite. $\sigma(x+3)$ reaches 0.5 at $x=-3$ (left). Rewrite it as $\sigma\big(w(x - x_0)\big)$ with $x_0 = -b/w$ to see the shift directly.

5The midpoint formula: where does the neuron switch?

The output is exactly 0.5 when $z=0$:

$$wx + b = 0 \;\Rightarrow\; x_0 = -\frac{b}{w} \qquad\text{so}\qquad \sigma(wx+b) = \sigma\big(w\,(x - x_0)\big)$$

This form separates the two jobs:

$x_0 = -b/w$: where

The location of the switch (decision boundary). Set by the ratio of $b$ to $w$.

$w$: how sharp

The steepness of the switch. Its sign sets the direction (rising or falling).

Worked example: $w=2,\ b=-6$ gives $x_0 = 3$, slope $2/4=0.5$ at $x=3$, rising. Check: $\sigma(2\cdot3-6)=\sigma(0)=0.5$ ✓.

6Training a sigmoid neuron is choosing w and b

Module 3.3 p. 20–26: fit $f(x)=\sigma(wx+b)$ to two points $(0.5,\,0.2)$ and $(2.5,\,0.9)$ with the loss

$$\mathcal{L}(w,b) = \tfrac12\Big[(0.9 - f(2.5))^2 + (0.2 - f(0.5))^2\Big] \quad\text{(p. 24)}$$

Click the snapshots from slide 26 to see how $w$ and $b$ move the curve onto the points. Or drag the sliders and try to beat the loss yourself.

f(0.5) → 0.2
f(2.5) → 0.9
Loss

Read the snapshots from top to bottom: w grows (the curve gets steeper, so it can climb from 0.2 to 0.9 between x=0.5 and x=2.5) and b becomes more negative (the curve slides right so the 0.5 crossing falls between the points, at $x_0=2.27/1.78\approx1.28$). Gradient descent p. 41–46 automates exactly these nudges using

$$\nabla_w = (f(x)-y)\,f(x)(1-f(x))\,x, \qquad \nabla_b = (f(x)-y)\,f(x)(1-f(x)) \quad\text{(p. 44)}$$

Note the extra $x$ in $\nabla_w$: the weight's gradient scales with the input, but the bias's does not, because $\partial z/\partial b = 1$.

7Large w gives a step; two steps give a tower

Module 3.5 p. 56: “If we take the logistic function and set w to a very high value we will recover the step function.” Then $b$ chooses where the step sits ($x_0=-b/w$). Subtract two steps placed at different points and you get a tower p. 56–58:

$$h_{21} = h_{11} - h_{12} = \sigma(w x + b_1) - \sigma(w x + b_2), \qquad \text{tower from } -\tfrac{b_1}{w} \text{ to } -\tfrac{b_2}{w}$$
h₁₁ = σ(wx + b₁) h₁₂ = σ(wx + b₂) tower h₂₁ = h₁₁ − h₁₂
b₁ = −w·left
b₂ = −w·right

8Two inputs: $\sigma(w_1x_1 + w_2x_2 + b)$

With two inputs p. 59–64 the sigmoid becomes a 3D surface, and the same rules apply in each direction:

9Cheat sheet

ChangeWhat happens to σ(wx + b)
↑ |w|Steeper. Approaches a step function
↓ |w| toward 0Flatter. At $w=0$ it is a constant $\sigma(b)$
w → −wCurve flips (decreasing)
↑ bShifts left by $\Delta b/w$ (fires earlier)
↓ bShifts right (fires later, i.e. a higher threshold)
Scale w and b together by kSame midpoint $-b/w$, but k× steeper
Midpoint (output 0.5)$x_0=-b/w$
Slope at the midpoint$w/4$, the maximum slope
Output rangeAlways $(0,1)$. $w$ and $b$ never change the range
Exam traps: (1) $+b$ shifts left, not right. (2) Changing $w$ also moves the midpoint when $b\neq0$. (3) A negative $w$ flips the curve. It does not shift it. (4) Neither $w$ nor $b$ changes the output range (0, 1).

10Practice questions (click to reveal)

Where does $\sigma(4x - 8)$ cross 0.5, and what is its slope there?

$x_0 = 8/4 = 2$. The slope is $4/4 = 1$.

$\sigma(x)$ and $\sigma(x+5)$: which one is further left?

$\sigma(x+5)$. Its midpoint is at $x=-5$.

You want a neuron that switches sharply from 1 to 0 at $x=3$. Give one choice of w, b.

Use a negative, large $w$: e.g. $w=-20$, $b=60$, since $-b/w = 60/20 = 3$.

What is $\sigma(0\cdot x + b)$ for $b=0$? For $b$ very large?

Constant 0.5 for $b=0$. Constant ≈ 1 for very large $b$. In both cases the input has no effect.

σ(w x + b) with w = 2, b = −4 vs. w = 20, b = −40. What is the same and what is different?

Same midpoint $x_0=2$. The second is 10× steeper (slope 5 vs. 0.5), so it is almost a step.

In the p. 26 training snapshots, why does b become more negative as training proceeds?

The midpoint $-b/w$ must sit between $x=0.5$ and $x=2.5$. As $w$ grows to make the curve steeper, $b$ must become more negative to keep the midpoint around $x\approx1.3$.

Make a tower of height 1 between $x=1$ and $x=4$ using two sigmoids with $w=50$.

$\sigma(50x - 50) - \sigma(50x - 200)$, i.e. $b_1=-50$ and $b_2=-200$.