Sigmoid Spline Experiment

Active

A PyTorch experiment on vanishing gradients. I compare a spline-modified sigmoid with ReLU and sigmoid, exploring what happens when deep networks learn.

Started
2025-07-08
Updated
2025-09-15
pythonpytorchdeep-learningactivation-functionsplinevanishing-gradient

Background and motivation

During my studies in 2025, I wanted to explore the mathematical foundations of neural networks, especially the vanishing gradient problem. I investigated whether modifying the sigmoid activation could make it useful in a deep network.

The idea came from my mathematics lectures. I used cubic Hermite splines to change the saturation behaviour of sigmoid and investigated the effect on gradient flow.

An intentionally difficult setup

I used a setup designed to expose vanishing gradients and compare activation functions.

  • A deep feedforward MLP with 16 hidden layers of 512 neurons each.
  • No residual connections, batch normalisation or dropout, to isolate the role of the activation function.
  • 250 training epochs with Adam on an NVIDIA RTX 4070 Ti Super.

I compared ReLU, standard sigmoid and my optimised sigmoid spline.

The spline modification

The modified function approximates sigmoid with splines in the central region and continues linearly beyond it. This keeps the derivative from approaching zero in the outer regions, unlike the standard sigmoid.

Modified sigmoid curve and its derivative, including the linear continuation

The function on the left and its derivative on the right. Linear continuation keeps the derivative nonzero in the outer regions.

Results

The experiments on FashionMNIST, CIFAR-10 and Tiny ImageNet-200 showed different learning behaviour across the datasets.

FashionMNIST accuracy curves comparing ReLU, sigmoid and the spline variant

On FashionMNIST, the spline variant approaches ReLU after an initial phase, while standard sigmoid stays near chance level.

DatasetReLUSigmoidSpline
FashionMNIST90.1%10.0%86.9%
CIFAR-1053.9%10.0%13.1%
Tiny ImageNet8.4%0.5%8.9%

Looking at the gradients

With standard sigmoid, gradients fell to around 10⁻¹¹ in this deep network and learning stalled near chance level.

Heatmap showing very small gradients with standard sigmoid on FashionMNIST

The sigmoid heatmap shows the weak gradient signal.

In the investigated FashionMNIST runs, the spline variant preserved gradient flow compared with standard sigmoid, including in earlier layers.

Heatmap showing stronger gradients with the spline variant on FashionMNIST

Stronger gradient signals are visible with the spline variant.

Better gradient flow did not automatically lead to competitive classification accuracy. On CIFAR-10, for example, the spline model remained well below ReLU. The results need to be considered in the context of this simple MLP architecture and the individual datasets.

Conclusions and runtime

In this setup, the spline modification improved gradient flow compared with standard sigmoid. This does not establish a general solution to the vanishing gradient problem. Classification performance depended strongly on the dataset, and the additional computation also matters.

Training time comparison for ReLU, sigmoid and the spline variant

The spline variant required more computation than ReLU in these runs.

The most useful lesson for me was to look beyond a single accuracy number. Learning curves, gradients and runtime together show what an approach achieves in these experiments and where its limits are.