Classifying bird species from sound
The project focuses on selecting a model and evaluation criterion that remain informative when the class distribution is imbalanced.
Explore one multi-input architecture.
An experimental comparison of model families, hyperparameters, acoustic representations, and evaluation criteria for bird-species recognition.
Mel maps
Three Mel representations pass through a two-dimensional convolutional stack before being flattened to 192 features.
Each input follows a model suited to its shape.
Mel and MFCC inputs use separate 2D CNNs, while the engineered feature sequence uses a 1D CNN. Their outputs are flattened, concatenated, and passed to a seven-class classifier.

Every species gets equal weight.
Macro-F1 calculates a score per class before averaging, so frequent species cannot hide weak performance on rarer classes.

Which model type and hyperparameter configuration performs best for bird-call classification, and which evaluation criterion gives the most reliable comparison under a strongly imbalanced class distribution?
- 01
Prepared Mel spectrograms, MFCCs with deltas, and engineered acoustic descriptors from each recording.
- 02
Compared feed-forward, convolutional, recurrent, and multi-input architectures across different hyperparameter configurations using cross-validation.
- 03
Tested specialized branches that process each representation according to its tensor structure before feature fusion.
- 04
Selected macro-F1 as the primary criterion so frequent species could not dominate the model comparison.
- Chose
- Made macro-F1 the primary selection criterion
- Instead of
- Ranking models by overall accuracy
- Because
- The class distribution is strongly imbalanced. Accuracy lets frequent species mask weak performance on rare ones, which is exactly the failure the model comparison needed to expose.
- PyTorch
- 2D CNN
- 1D CNN
- Recurrent models
- MFCC · Mel
- Cross-validation
- Macro-F1