Vanilla Neural Networks

A neural network is just a math function: y = FNN(x).

FNN has a nested form. Think of it as a stack of layers. A 3-layer neural network that returns a scalar value looks like this:

y = FNN(x) = f₃(f₂(f₁(x)))

x flows through three layers, f one, f two, f three, each computing g of W x plus b, to produce y.

Each f — f₁, f₂, … fₙ — has the same form:

f(x) = g(Wx + b)

W (the weight matrix) and b (a bias vector) are the learned parameters, usually trained via gradient descent. g is the activation function, and it can be chosen differently for each layer.

Wx + b is linear — wrapping it in g is what makes each layer non-linear. Without g (or with g chosen to be linear), the whole FNN collapses into a single linear function: stack 100 such layers and the composition of linear maps is still just one linear map, no matter how deep the network looks. g is what lets a stack of layers approximate anything more than a straight line. Popular choices for g are sigmoid and ReLU.

There are many variants of neural networks — CNNs, RNNs, transformers, and more — each shaped by assumptions about the data they process. The example above, where every neuron in one layer connects to every neuron in the next, is the plainest of them: a multilayer perceptron (MLP), also called a vanilla neural network.