Sigmoid Spline Experiment
ActiveA PyTorch experiment on vanishing gradients. I compare a spline-modified sigmoid with ReLU and sigmoid, exploring what happens when deep networks learn.
- Started
- 2025-07-08
- Updated
- 2025-09-15
Background and motivation
During my studies in 2025, I wanted to explore the mathematical foundations of neural networks, especially the vanishing gradient problem. I investigated whether modifying the sigmoid activation could make it useful in a deep network.
The idea came from my mathematics lectures. I used cubic Hermite splines to change the saturation behaviour of sigmoid and investigated the effect on gradient flow.
An intentionally difficult setup
I used a setup designed to expose vanishing gradients and compare activation functions.
- A deep feedforward MLP with 16 hidden layers of 512 neurons each.
- No residual connections, batch normalisation or dropout, to isolate the role of the activation function.
- 250 training epochs with Adam on an NVIDIA RTX 4070 Ti Super.
I compared ReLU, standard sigmoid and my optimised sigmoid spline.
The spline modification
The modified function approximates sigmoid with splines in the central region and continues linearly beyond it. This keeps the derivative from approaching zero in the outer regions, unlike the standard sigmoid.

The function on the left and its derivative on the right. Linear continuation keeps the derivative nonzero in the outer regions.
Results
The experiments on FashionMNIST, CIFAR-10 and Tiny ImageNet-200 showed different learning behaviour across the datasets.

On FashionMNIST, the spline variant approaches ReLU after an initial phase, while standard sigmoid stays near chance level.
| Dataset | ReLU | Sigmoid | Spline |
|---|---|---|---|
| FashionMNIST | 90.1% | 10.0% | 86.9% |
| CIFAR-10 | 53.9% | 10.0% | 13.1% |
| Tiny ImageNet | 8.4% | 0.5% | 8.9% |
Looking at the gradients
With standard sigmoid, gradients fell to around 10⁻¹¹ in this deep network and learning stalled near chance level.

The sigmoid heatmap shows the weak gradient signal.
In the investigated FashionMNIST runs, the spline variant preserved gradient flow compared with standard sigmoid, including in earlier layers.

Stronger gradient signals are visible with the spline variant.
Better gradient flow did not automatically lead to competitive classification accuracy. On CIFAR-10, for example, the spline model remained well below ReLU. The results need to be considered in the context of this simple MLP architecture and the individual datasets.
Conclusions and runtime
In this setup, the spline modification improved gradient flow compared with standard sigmoid. This does not establish a general solution to the vanishing gradient problem. Classification performance depended strongly on the dataset, and the additional computation also matters.

The spline variant required more computation than ReLU in these runs.
The most useful lesson for me was to look beyond a single accuracy number. Learning curves, gradients and runtime together show what an approach achieves in these experiments and where its limits are.