PROJECT / 04Audio ML · Feature fusion

Classifying bird species from sound

The project focuses on selecting a model and evaluation criterion that remain informative when the class distribution is imbalanced.

Role
Neural-network design and evaluation
Context
Bird-species audio classification
Period
Academic project
INTERACTIVE EVIDENCE

Explore one multi-input architecture.

An experimental comparison of model families, hyperparameters, acoustic representations, and evaluation criteria for bird-species recognition.

MULTI-INPUT CNN / SELECT A BRANCH Mel · MFCC · engineered features
CONCATENATE1,381 combined features192 + 1,125 + 64
SHARED CLASSIFIERFeed-forward networkseven species outputs
SELECTED / 2D CNN

Mel maps

Three Mel representations pass through a two-dimensional convolutional stack before being flattened to 192 features.

DETAILED MODEL ARCHITECTURE

Each input follows a model suited to its shape.

Mel and MFCC inputs use separate 2D CNNs, while the engineered feature sequence uses a 1D CNN. Their outputs are flattened, concatenated, and passed to a seven-class classifier.

Original bird-classification model graph with separate Mel, MFCC, and engineered-feature convolutional branches
WHY MACRO-F1

Every species gets equal weight.

Macro-F1 calculates a score per class before averaging, so frequent species cannot hide weak performance on rarer classes.

Recorded MNN training and validation macro-F1 curves
THE QUESTION

Which model type and hyperparameter configuration performs best for bird-call classification, and which evaluation criterion gives the most reliable comparison under a strongly imbalanced class distribution?

HOW I APPROACHED IT
  1. 01

    Prepared Mel spectrograms, MFCCs with deltas, and engineered acoustic descriptors from each recording.

  2. 02

    Compared feed-forward, convolutional, recurrent, and multi-input architectures across different hyperparameter configurations using cross-validation.

  3. 03

    Tested specialized branches that process each representation according to its tensor structure before feature fusion.

  4. 04

    Selected macro-F1 as the primary criterion so frequent species could not dominate the model comparison.

The call I made
Chose
Made macro-F1 the primary selection criterion
Instead of
Ranking models by overall accuracy
Because
The class distribution is strongly imbalanced. Accuracy lets frequent species mask weak performance on rare ones, which is exactly the failure the model comparison needed to expose.
Model familiesFFN · CNN · recurrent
RepresentationsMel · MFCC · 55 features
Primary metricMacro-F1
TOOLS & METHODS
  • PyTorch
  • 2D CNN
  • 1D CNN
  • Recurrent models
  • MFCC · Mel
  • Cross-validation
  • Macro-F1
Inspect the source repository