← All repos

Dynamic Contextual Emotions Transformer (DCET)

Multi-stage deep learning framework for temporal emotion recognition in conversations. Tracks emotional shifts across dialogue sequences using visual, audio, and text modalities with Recurrent Memory Transformer. Modular codebase with clean package structure, YAML-based config management, and configurable paths. Stage 0 (preprocessing) extracts features from video. Stage 1 (multimodal fusion) combines frozen ResNet-50 + Wav2Vec2 backbones. Stage 2 (temporal modeling) processes utterance sequences with RMT for per-speaker trajectory tracking. Production-ready with reproducible results tracking via date-based naming conventions.

README.md

Dynamic Contextual Emotions Transformer (DCET)

1. Project Overview

This project presents the Dynamic Contextual Emotions Transformer (DCET), a novel multi-stage cascaded deep learning framework designed for robust, real-time temporal tracking of emotional shifts in conversational settings. Addressing the limitations of existing multimodal models that often struggle to adapt to the dynamic nature of real-world interactions, DCET introduces an Adaptive Contextual Fusion mechanism. Our architecture not only identifies static emotional states but also adaptively weighs and integrates information from visual, audio, and textual modalities based on their real-time informativeness within the evolving conversational context. This is achieved through a combination of specialized unimodal feature extractors, cross-modal attention for dynamic fusion, and a recurrent memory network for maintaining long-term, per-speaker emotional trajectories. The ultimate goal is to pave the way for more empathetic and effective conversational agents capable of truly understanding and responding to human emotion in live, interactive environments like Augmented Reality (AR).

2. Key Features

  • ✨ Multi-stage Architecture: Clear separation of preprocessing, fusion, and temporal modeling
  • 🧠 Recurrent Memory Transformer (RMT): Learnable memory for tracking dialogue context
  • šŸ‘¤ Per-Speaker Tracking: Individual emotional trajectories across conversations
  • šŸš€ Production-Ready: Organized codebase with YAML configs, no hardcoded paths
  • šŸ“Š Reproducible: All hyperparameters in config files, metadata tracking in results
  • šŸ”§ Extensible: Easy to add new models, datasets, or preprocessing steps

3. Repository Structure

/Dyn/
ā”œā”€ā”€ src/                           # Main package
│   ā”œā”€ā”€ models/                    # Stage 1 & 2 model definitions
│   ā”œā”€ā”€ datasets/                  # Data loading utilities
│   ā”œā”€ā”€ preprocessing/             # Feature extraction
│   ā”œā”€ā”€ training/                  # Training utilities (templates)
│   └── utils/                     # Config + metrics
│
ā”œā”€ā”€ scripts/                       # Entry point scripts
│   ā”œā”€ā”€ preprocess.py             # Stage 0: Feature extraction
│   ā”œā”€ā”€ train_stage1.py           # Stage 1: Multimodal fusion training
│   ā”œā”€ā”€ train_stage2.py           # Stage 2: Temporal model training
│   └── evaluate.py               # Inference & evaluation
│
ā”œā”€ā”€ configs/                       # YAML configuration files
│   ā”œā”€ā”€ stage1_config.yaml
│   └── stage2_config.yaml
│
ā”œā”€ā”€ data/                          # Data directory structure
│   ā”œā”€ā”€ raw/                       # Raw videos + annotations
│   ā”œā”€ā”€ preprocessed/              # Extracted .npy features
│   └── splits/                    # Train/val/test splits
│
ā”œā”€ā”€ results/                       # Training runs & checkpoints
│   ā”œā”€ā”€ stage1_runs/
│   └── stage2_runs/
│
ā”œā”€ā”€ .archive/original_code/        # Legacy code (reference only)
ā”œā”€ā”€ setup.py                       # Package installation
└── ORGANIZATION.md                # Full directory guide

4.1 Quick Start

4.1.1 Installation

# Clone repository
cd dyn

# Create virtual environment
python3 -m venv venv
source venv/bin/activate  # On Windows: venv\Scripts\activate

# Install package
pip install -e .

# Verify installation
python -c "import src; print('āœ… Installation successful')"

4.1.2 Preprocessing (Stage 0)

Extract visual and audio features from raw video files:

python scripts/preprocess.py \
    --video-dir data/raw \
    --annotations-csv data/raw/annotations.csv \
    --output-dir data/preprocessed \
    --limit 100  # Optional: limit for testing

Input: MP4 videos in data/raw/
Output: .npy files in data/preprocessed/{masked_faces,audio_chunks}/

4.1.3 Stage 1 Training (Multimodal Fusion)

Train fusion of visual and audio features:

python scripts/train_stage1.py --config configs/stage1_config.yaml

Customize via configs/stage1_config.yaml:

model:
  vision_backbone: "resnet50"
  audio_backbone: "wav2vec2_base"
  fusion_output_dim: 512

training:
  batch_size: 32
  learning_rate: 0.001
  num_epochs: 100

Output: Utterance embeddings in results/stage1_runs/

4.1.4 Stage 2 Training (Temporal Modeling)

Train temporal model on dialogue sequences:

python scripts/train_stage2.py --config configs/stage2_config.yaml

Customize via configs/stage2_config.yaml:

model:
  model_type: "RMT"  # or "TransformerXL_like", "TCN"
  hidden_dim: 256
  num_layers: 3

training:
  batch_size: 16
  num_epochs: 50

Output: Checkpoints & metrics in results/stage2_runs/

4.5 Evaluation & Inference

Run inference on test data:

python scripts/evaluate.py \
    --checkpoint results/stage2_runs/20251211_062006_rmt/checkpoint_best.pth \
    --data-dir data/preprocessed \
    --stage 2

5. Configuration

All training parameters are in YAML files, no hardcoded paths.

Stage 1 Config (configs/stage1_config.yaml)

data:
  preprocessed_data_dir: "data/preprocessed"
  annotations_csv: "data/raw/annotations.csv"

model:
  vision_feature_dim: 2048
  audio_feature_dim: 768
  fusion_output_dim: 512
  num_emotions: 7

training:
  batch_size: 32
  learning_rate: 0.001
  num_epochs: 100

Stage 2 Config (configs/stage2_config.yaml)

model:
  model_type: "RMT"           # "RMT", "TransformerXL_like", "TCN"
  embedding_dim: 512
  hidden_dim: 256
  num_layers: 3
  num_heads: 4

training:
  batch_size: 16
  learning_rate: 0.0005
  num_epochs: 50

6. Architecture Details

Stage 0: Preprocessing

  • Input: Raw MP4 video files
  • Process: Detect speaker face, extract audio samples
  • Output: .npy files with masked frames and audio

Stage 1: Multimodal Fusion

  • Vision: ResNet-50 (frozen) → aggregated features
  • Audio: Wav2Vec2-base (frozen) → aggregated features
  • Fusion: Concatenate → MLP (trainable)
  • Output: 512-dim utterance embeddings

Stage 2: Temporal Modeling

  • Input: Sequence of Stage 1 embeddings
  • Speaker Embeddings: Learned 64-dim vectors per speaker
  • Temporal Core Options:
    • RMT: Recurrent Memory Transformer with learnable memory
    • TransformerXL: Standard transformer with external state
    • TCN: Temporal Convolutional Network
  • Output: Per-utterance emotion + sentiment predictions

7. Results

Current Experiments

StageModelDateAccuracyStatus
1ResNet18 + Wav2Vec22025-02-01~45%āœ… Done
2RMT2025-12-11~88%āœ… Done

See results/README.md and results/ORGANIZATION_GUIDE.md for full details on result naming and structure.

Viewing Results

# List all runs (newest first)
ls -lt results/stage2_runs/

# View run metadata
cat results/stage2_runs/20251211_062006_rmt/README.md

# View metrics
cat results/stage2_runs/20251211_062006_rmt/metrics_stage2.csv

# View confusion matrix
open results/stage2_runs/20251211_062006_rmt/confusion_matrix_epoch_128.png

8. Dataset

The code expects data in this format:

data/
ā”œā”€ā”€ raw/
│   ā”œā”€ā”€ annotations.csv           # Columns: Dialogue_ID, Utterance_ID, Speaker, Emotion, Sentiment
│   └── *.mp4                     # Video files (format: dia{id}_utt{id}.mp4)
│
└── preprocessed/
    ā”œā”€ā”€ masked_faces/
    │   └── dia{id}_utt{id}_frames.npy
    └── audio_chunks/
        └── dia{id}_utt{id}_audio.npy

Annotation CSV Format

Dialogue_ID,Utterance_ID,Speaker,Emotion,Sentiment
1,1,Alice,neutral,positive
1,2,Bob,joy,positive
1,3,Alice,surprise,neutral

9. Usage Examples

Training Full Pipeline

# 1. Preprocess
python scripts/preprocess.py --video-dir data/raw --output-dir data/preprocessed

# 2. Train Stage 1
python scripts/train_stage1.py --config configs/stage1_config.yaml

# 3. Train Stage 2
python scripts/train_stage2.py --config configs/stage2_config.yaml

# 4. Evaluate
python scripts/evaluate.py --checkpoint results/stage2_runs/*/checkpoint_best.pth

Custom Configuration

# Edit config
nano configs/stage2_config.yaml

# Run with custom config
python scripts/train_stage2.py --config configs/stage2_config.yaml

Resume Training

python scripts/train_stage2.py \
    --config configs/stage2_config.yaml \
    --resume-from results/stage2_runs/20251211_062006_rmt/checkpoint_best.pth

10. Project Organization

For detailed documentation, see:

  • ORGANIZATION.md - Complete directory structure and workflow
  • FLAW_ANALYSIS.md - Analysis of structural issues (fixed)
  • REFACTOR_SUMMARY.md - Before/after comparison
  • results/README.md - Results structure and naming convention
  • results/ORGANIZATION_GUIDE.md - Best practices for result tracking

11. Development

Adding a New Model

  1. Create model class in src/models/
  2. Import in src/models/__init__.py
  3. Update training script template
  4. Test in config before training

Running Tests

python -m pytest tests/

Code Style

# Format code
black src/ scripts/

# Type checking
mypy src/

12. Important Notes

No Hardcoded Paths

All paths are relative and configurable via YAML:

from src.utils import load_config
config = load_config("configs/stage1_config.yaml")
data_dir = config["data"]["preprocessed_data_dir"]

Results Naming

Runs use date-based naming: YYYYMMDD_HHMMSS_model_type

  • āœ… Reproducible (no hardcoded metrics)
  • āœ… Sortable (chronological order)
  • āœ… Self-documenting (model visible)

Example: 20251211_062006_rmt = 2025-12-11 at 06:20:06, RMT model

Reproducibility

Each run should include:

  • Copy of config file used
  • metadata.json with hyperparameters
  • Random seeds logged
  • Git commit hash

13. License

GNU General Public License v3.0 (see LICENSE)

15. Support & Contributing

For issues, questions, or contributions, please refer to the developer documentation in ORGANIZATION.md.


Version: 0.1.0