article
A Vanilla GAN on MNIST, and What FID Actually Told Me
Generative Adversarial Networks pit two networks against each other: a generator trying to produce convincing fakes, a discriminator trying to catch them. Deliberately skipping convolutions for this one — no DCGAN, just fully-connected layers on both sides — to see how far the plain 2014 formulation gets on MNIST digits before reaching for anything fancier.
The architecture
Generator: 100 → 256 → 512 → 1024 → 784, ReLU between layers, Dropout(0.3), Tanh on the output to land in [-1, 1]. Discriminator: 784 → 512 → 256 → 1, LeakyReLU(0.2), Dropout(0.3), Sigmoid at the end. Both trained with Adam (lr=0.0002, betas=(0.5, 0.999)) and BCELoss, batch size 128, for 500 epochs — the standard DCGAN-paper hyperparameters, just without the convolutions.
Measuring it with FID, not vibes
The discriminator’s loss alone doesn’t tell you much about image quality — it only says whether the discriminator can currently tell real from fake, which oscillates by design. Fréchet Inception Distance instead compares the actual distributions of real and generated images (via a pretrained Inception network’s features), so I logged it every epoch alongside the losses.
Where it got interesting
Computing FID per mini-batch on only 128 images turned out to be its own experiment in numerical instability. Covariance matrices estimated from that few samples are badly conditioned, and the matrix square root in the FID formula would occasionally explode:
2.83e+32
-8.85e+39
-1.96e+58
Values like that aren’t a signal, they’re sqrtm failing quietly on an ill-conditioned matrix. The fix was a straightforward post-hoc filter — keep anything in a sane range, throw out anything outside (-1e6, 1e6) — before plotting. That’s not a workaround I’d trust in a paper, but for tracking training trend on a side project it’s exactly the pragmatic level of rigor the situation called for.
What the filtered numbers showed
Once filtered, FID dropped quickly from a high starting value to under 100 within the first several dozen epochs and stayed roughly flat after that — the generator found a stable regime early and didn’t wander off it. The loss curves back that up: discriminator loss hovered close to 1 for the whole run (still finding real-vs-fake genuinely hard, which is the healthy outcome), and generator loss dropped fast and then sat at a low, stable level. No mode collapse, no runaway divergence between the two networks — a boringly well-behaved 500-epoch GAN, which for a first fully-connected GAN is exactly the result you want.
Source: github.com/chris017/GAN