Dynamic Contextual Emotions Transformer (DCET)
1. Project Overview
This project presents the Dynamic Contextual Emotions Transformer (DCET), a novel multi-stage cascaded deep learning framework designed for robust, real-time temporal tracking of emotional shifts in conversational settings. Addressing the limitations of existing multimodal models that often struggle to adapt to the dynamic nature of real-world interactions, DCET introduces an Adaptive Contextual Fusion mechanism. Our architecture not only identifies static emotional states but also adaptively weighs and integrates information from visual, audio, and textual modalities based on their real-time informativeness within the evolving conversational context. This is achieved through a combination of specialized unimodal feature extractors, cross-modal attention for dynamic fusion, and a recurrent memory network for maintaining long-term, per-speaker emotional trajectories. The ultimate goal is to pave the way for more empathetic and effective conversational agents capable of truly understanding and responding to human emotion in live, interactive environments like Augmented Reality (AR).
2. Key Features
- ⨠Multi-stage Architecture: Clear separation of preprocessing, fusion, and temporal modeling
- š§ Recurrent Memory Transformer (RMT): Learnable memory for tracking dialogue context
- š¤ Per-Speaker Tracking: Individual emotional trajectories across conversations
- š Production-Ready: Organized codebase with YAML configs, no hardcoded paths
- š Reproducible: All hyperparameters in config files, metadata tracking in results
- š§ Extensible: Easy to add new models, datasets, or preprocessing steps
3. Repository Structure
/Dyn/
āāā src/ # Main package
ā āāā models/ # Stage 1 & 2 model definitions
ā āāā datasets/ # Data loading utilities
ā āāā preprocessing/ # Feature extraction
ā āāā training/ # Training utilities (templates)
ā āāā utils/ # Config + metrics
ā
āāā scripts/ # Entry point scripts
ā āāā preprocess.py # Stage 0: Feature extraction
ā āāā train_stage1.py # Stage 1: Multimodal fusion training
ā āāā train_stage2.py # Stage 2: Temporal model training
ā āāā evaluate.py # Inference & evaluation
ā
āāā configs/ # YAML configuration files
ā āāā stage1_config.yaml
ā āāā stage2_config.yaml
ā
āāā data/ # Data directory structure
ā āāā raw/ # Raw videos + annotations
ā āāā preprocessed/ # Extracted .npy features
ā āāā splits/ # Train/val/test splits
ā
āāā results/ # Training runs & checkpoints
ā āāā stage1_runs/
ā āāā stage2_runs/
ā
āāā .archive/original_code/ # Legacy code (reference only)
āāā setup.py # Package installation
āāā ORGANIZATION.md # Full directory guide
4.1 Quick Start
4.1.1 Installation
# Clone repository
cd dyn
# Create virtual environment
python3 -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
# Install package
pip install -e .
# Verify installation
python -c "import src; print('ā
Installation successful')"
4.1.2 Preprocessing (Stage 0)
Extract visual and audio features from raw video files:
python scripts/preprocess.py \
--video-dir data/raw \
--annotations-csv data/raw/annotations.csv \
--output-dir data/preprocessed \
--limit 100 # Optional: limit for testing
Input: MP4 videos in data/raw/
Output: .npy files in data/preprocessed/{masked_faces,audio_chunks}/
4.1.3 Stage 1 Training (Multimodal Fusion)
Train fusion of visual and audio features:
python scripts/train_stage1.py --config configs/stage1_config.yaml
Customize via configs/stage1_config.yaml:
model:
vision_backbone: "resnet50"
audio_backbone: "wav2vec2_base"
fusion_output_dim: 512
training:
batch_size: 32
learning_rate: 0.001
num_epochs: 100
Output: Utterance embeddings in results/stage1_runs/
4.1.4 Stage 2 Training (Temporal Modeling)
Train temporal model on dialogue sequences:
python scripts/train_stage2.py --config configs/stage2_config.yaml
Customize via configs/stage2_config.yaml:
model:
model_type: "RMT" # or "TransformerXL_like", "TCN"
hidden_dim: 256
num_layers: 3
training:
batch_size: 16
num_epochs: 50
Output: Checkpoints & metrics in results/stage2_runs/
4.5 Evaluation & Inference
Run inference on test data:
python scripts/evaluate.py \
--checkpoint results/stage2_runs/20251211_062006_rmt/checkpoint_best.pth \
--data-dir data/preprocessed \
--stage 2
5. Configuration
All training parameters are in YAML files, no hardcoded paths.
Stage 1 Config (configs/stage1_config.yaml)
data:
preprocessed_data_dir: "data/preprocessed"
annotations_csv: "data/raw/annotations.csv"
model:
vision_feature_dim: 2048
audio_feature_dim: 768
fusion_output_dim: 512
num_emotions: 7
training:
batch_size: 32
learning_rate: 0.001
num_epochs: 100
Stage 2 Config (configs/stage2_config.yaml)
model:
model_type: "RMT" # "RMT", "TransformerXL_like", "TCN"
embedding_dim: 512
hidden_dim: 256
num_layers: 3
num_heads: 4
training:
batch_size: 16
learning_rate: 0.0005
num_epochs: 50
6. Architecture Details
Stage 0: Preprocessing
- Input: Raw MP4 video files
- Process: Detect speaker face, extract audio samples
- Output:
.npyfiles with masked frames and audio
Stage 1: Multimodal Fusion
- Vision: ResNet-50 (frozen) ā aggregated features
- Audio: Wav2Vec2-base (frozen) ā aggregated features
- Fusion: Concatenate ā MLP (trainable)
- Output: 512-dim utterance embeddings
Stage 2: Temporal Modeling
- Input: Sequence of Stage 1 embeddings
- Speaker Embeddings: Learned 64-dim vectors per speaker
- Temporal Core Options:
- RMT: Recurrent Memory Transformer with learnable memory
- TransformerXL: Standard transformer with external state
- TCN: Temporal Convolutional Network
- Output: Per-utterance emotion + sentiment predictions
7. Results
Current Experiments
| Stage | Model | Date | Accuracy | Status |
|---|---|---|---|---|
| 1 | ResNet18 + Wav2Vec2 | 2025-02-01 | ~45% | ā Done |
| 2 | RMT | 2025-12-11 | ~88% | ā Done |
See results/README.md and results/ORGANIZATION_GUIDE.md for full details on result naming and structure.
Viewing Results
# List all runs (newest first)
ls -lt results/stage2_runs/
# View run metadata
cat results/stage2_runs/20251211_062006_rmt/README.md
# View metrics
cat results/stage2_runs/20251211_062006_rmt/metrics_stage2.csv
# View confusion matrix
open results/stage2_runs/20251211_062006_rmt/confusion_matrix_epoch_128.png
8. Dataset
The code expects data in this format:
data/
āāā raw/
ā āāā annotations.csv # Columns: Dialogue_ID, Utterance_ID, Speaker, Emotion, Sentiment
ā āāā *.mp4 # Video files (format: dia{id}_utt{id}.mp4)
ā
āāā preprocessed/
āāā masked_faces/
ā āāā dia{id}_utt{id}_frames.npy
āāā audio_chunks/
āāā dia{id}_utt{id}_audio.npy
Annotation CSV Format
Dialogue_ID,Utterance_ID,Speaker,Emotion,Sentiment
1,1,Alice,neutral,positive
1,2,Bob,joy,positive
1,3,Alice,surprise,neutral
9. Usage Examples
Training Full Pipeline
# 1. Preprocess
python scripts/preprocess.py --video-dir data/raw --output-dir data/preprocessed
# 2. Train Stage 1
python scripts/train_stage1.py --config configs/stage1_config.yaml
# 3. Train Stage 2
python scripts/train_stage2.py --config configs/stage2_config.yaml
# 4. Evaluate
python scripts/evaluate.py --checkpoint results/stage2_runs/*/checkpoint_best.pth
Custom Configuration
# Edit config
nano configs/stage2_config.yaml
# Run with custom config
python scripts/train_stage2.py --config configs/stage2_config.yaml
Resume Training
python scripts/train_stage2.py \
--config configs/stage2_config.yaml \
--resume-from results/stage2_runs/20251211_062006_rmt/checkpoint_best.pth
10. Project Organization
For detailed documentation, see:
ORGANIZATION.md- Complete directory structure and workflowFLAW_ANALYSIS.md- Analysis of structural issues (fixed)REFACTOR_SUMMARY.md- Before/after comparisonresults/README.md- Results structure and naming conventionresults/ORGANIZATION_GUIDE.md- Best practices for result tracking
11. Development
Adding a New Model
- Create model class in
src/models/ - Import in
src/models/__init__.py - Update training script template
- Test in config before training
Running Tests
python -m pytest tests/
Code Style
# Format code
black src/ scripts/
# Type checking
mypy src/
12. Important Notes
No Hardcoded Paths
All paths are relative and configurable via YAML:
from src.utils import load_config
config = load_config("configs/stage1_config.yaml")
data_dir = config["data"]["preprocessed_data_dir"]
Results Naming
Runs use date-based naming: YYYYMMDD_HHMMSS_model_type
- ā Reproducible (no hardcoded metrics)
- ā Sortable (chronological order)
- ā Self-documenting (model visible)
Example: 20251211_062006_rmt = 2025-12-11 at 06:20:06, RMT model
Reproducibility
Each run should include:
- Copy of config file used
metadata.jsonwith hyperparameters- Random seeds logged
- Git commit hash
13. License
GNU General Public License v3.0 (see LICENSE)
15. Support & Contributing
For issues, questions, or contributions, please refer to the developer documentation in ORGANIZATION.md.
Version: 0.1.0