โ† All repos

Open ML Foundry

Open-source multi-modal fine-tuning. LLMs and vision models, fully local, accelerate edge deployments. All-in-one framework for fine-tuning large language models and vision models on your machine. Session-based training for Qwen3.8, GLM-5.3-Flash, Kimi K3, MiniMax-H3, DeepSeek-V4, Gemma 4 with LoRA/QLoRA. Plus vision models (ResNet, YOLO, CLIP). Privacy-first, runs completely local, deploy to edge devices. No cloud required.

README.md
<p align="center"> <img src="assets/icon.ico" alt="Open ML Foundry" width="520"> </p>

Open-source multi-modal fine-tuning. LLMs and vision models, fully local, accelerate edge deployments

All-in-one framework for fine-tuning large language models and vision models on your machine. Session-based training for Qwen3.8, GLM-5.3-Flash, Kimi K3, MiniMax-H3, DeepSeek-V4, Gemma 4 with LoRA/QLoRA. Plus vision models (ResNet, YOLO, CLIP). Privacy-first, runs completely local, deploy to edge devices. No cloud required.

Python 3.8+ JAX/Flax Alpha


๐ŸŽฏ Use Cases (What You Can Do Today)

LLM Fine-Tuning (Session-Based) - v0.3.0

Use CaseFeatureModels
Fine-tune LLMLoRA/QLoRA with session historyQwen, Gemma, DeepSeek, GLM, MiniMax
Test inferenceChat-like interface during trainingAll LLM models
Model selectionChoose from HF Hub or GGUF modelsHF Hub + GGUF
Multi-turn trainingTrain on dialogue/conversation dataAll LLM models

See Known Limitations โ€” this hasn't been run end-to-end yet.

Vision Fine-Tuning

Use CaseFeatureStatus
Custom image classifierResNet/CNN + your data (demo-scale training loop)โœ… Working
Object detectionFasterRCNN ResNet50 (COCO)โœ… Working
Mobile exportTFLite/ONNX + quantizationโœ… Working
Edge inference8-10ms latencyโœ… Working
Image searchCLIP embeddingsโœ… Working
Model versioningAuto-versioning + rollbackโœ… Working
Test a fine-tuned checkpointLoad checkpoint โ†’ predictโœ… Working (FinetuneInference loads real params and runs a real forward pass)

๐Ÿ“Š Comparison with Existing Tools

LLM Fine-Tuning Focus

FeatureOpen ML FoundryUnslothLLaMA-FactoryAxolotl
LoRA/QLoRAโœ…โœ…โญ (fastest)โœ…โœ…
Session-based UIโœ… Chat historyโš ๏ธ Unsloth Studio (web UI + desktop app, has session features)โŒ CLI/Web basicโŒ CLI only
Model switchingโŒ New session per model (model_name is fixed at session creation, no code path changes it)โŒโš ๏ธ Restartโš ๏ธ Restart
Local-only (no cloud)โœ…โœ…โœ…โœ…
Popular modelsQwen, Gemma, DeepSeek, GLMโญ OptimizedAll HF modelsAll HF models
Inference UIโœ… Chat interfaceโš ๏ธ Unsloth Studio (chat + side-by-side model comparison)โš ๏ธ BasicโŒ
Vision + LLMโœ… BothโŒ LLM onlyโŒ LLM onlyโŒ LLM only

Vision Fine-Tuning (Legacy Support)

FeatureOpen ML FoundryPyTorch LightningFastAI
Image classificationโœ… ResNet/CNNโœ…โœ…
Object detectionโœ… FasterRCNNโŒโš ๏ธ
Edge deploymentโœ… TFLite/ONNXโš ๏ธโŒ
CLIP embeddingsโœ…โŒโŒ

TL;DR - Choose based on use case:

  • Unsloth if: You want FASTEST LLM training speed (2-5x)
  • Open ML Foundry if: You need LLM + vision + edge deployment + session-based UI
  • LLaMA-Factory if: You want all HuggingFace models with advanced config
  • PyTorch Lightning if: You need distributed/multi-GPU framework

๐Ÿ“– Documentation (Streamlined for Users)

Start experimenting with built-in models immediately โ€” no setup needed beyond pip install and docker-compose up.

Essential Guides

  1. Models & Sample Experiments โ† Start here

    • Object detection on any image
    • Train custom classifier in 5-10s
    • Semantic image search
    • 5 quick experiment ideas
  2. CLI Guide โ† Command-line interface (Phase 1)

    • sentinel model import - Import custom models
    • sentinel dataset prepare - Prepare training data
    • sentinel train start - Train with built-in or custom models
    • Usage examples for all commands
  3. Getting Started (5 mins)

    • Install โ†’ Start โ†’ Use
  4. API Reference

    • All endpoints
    • Request/response examples
  5. Architecture Overview

    • How the system works
  6. Training Features

    • Fine-tuning capabilities
    • Model versioning
    • Performance metrics
  7. Inference Optimization

    • XLA speedups
    • Mobile export
    • Benchmark results
  8. Security & Auth

    • JWT tokens
    • Role-based access
    • Production hardening

๐Ÿš€ Get Started (5 Minutes)

1. Clone & Setup

git clone https://github.com/agentic-inquisit/open-ml-foundry.git
cd open-ml-foundry

# Run setup (checks dependencies, creates venv, installs packages)
./setup.sh

2. Session-Based Fine-Tuning (LLM + Vision, Chat UI)

# Start the API + session UI
uvicorn serving.main:app --reload --port 8000

# Open http://localhost:8000/sessions in a browser:
# 1. "+ New Session" โ†’ pick LLM (Qwen, Gemma, DeepSeek, GLM, MiniMax) or Vision
# 2. Point at a dataset (JSONL for LLM, image path for vision)
# 3. Start Training โ†’ progress streams into the session as chat events
# 4. Test tab โ†’ send a prompt/image, see the model's reply in the same thread
# 5. Every session keeps its full history โ€” switch sessions from the sidebar

Or drive the same flow over the REST API directly โ€” see Core API Endpoints below.

3. Vision Fine-Tuning (Direct API, no session)

# Object detection (pretrained, no fine-tuning needed)
curl http://localhost:8001/detect -F "image=@test.jpg"

# One-shot fine-tune (legacy endpoint, still works outside sessions)
curl -X POST http://localhost:8001/finetune -F "dataset=@image.jpg" -F "target_object=my_class"

4. Full Server Setup

# Terminal 1: Start all services
./start.sh

# Terminal 2: Test LLM inference
curl -X POST http://localhost:8000/llm/inference \
  -H "Content-Type: application/json" \
  -d '{"session_id": "my-session", "prompt": "Hello, how are you?"}'

# Test vision inference
curl http://localhost:8001/detect -F "image=@test.jpg"

โšก Core Features

๐Ÿ’ฌ LLM Fine-Tuning (Session-Based) - v0.3.0

Implemented in core/session_store.py, llm/, serving/session_api.py, serving/static/sessions_chat.html. Not yet run end-to-end โ€” see Known Limitations.

  • Session management - Each fine-tuning is a session with persistent, chat-like history (SQLite)
  • LoRA/QLoRA - via peft + transformers; QLoRA needs bitsandbytes + CUDA GPU (not installed by default)
  • Model support - Qwen, Gemma, DeepSeek, GLM, MiniMax, Kimi (via HF Hub โ€” see llm/supported_models.py for repo id verification status)
  • Model formats - HuggingFace Hub + GGUF (GGUF needs llama-cpp-python, not installed by default โ€” build-tool dependency)
  • Inference in-session - Test the model mid-conversation from the same chat UI
  • Multi-turn support - JSONL datasets with messages (chat) or prompt/completion shape
  • Checkpointing - LoRA adapter saved per session under training_outputs/llm/<session_id>/

๐Ÿ–ผ๏ธ Vision Fine-Tuning (Maintained)

  • Image classification - Fine-tune ResNet/CNN on your images
  • Object detection - FasterRCNN ResNet50+FPN (COCO pretrained)
  • Quick training - 5-10s for demo datasets, scales to hours
  • Validation & early stopping - Automatic train/val split with patience
  • Auto-versioning - v1.0, v1.1, v2.0 format with rollback

๐Ÿ“Š Model Management

  • Performance tracking - Loss, accuracy, metrics per epoch in SQLite
  • Model comparison - Compare versions side-by-side
  • Session history - View all training runs like conversations
  • Checkpoint storage - Save/restore at any epoch

๐Ÿš€ Optimization & Export

  • LoRA adapter export - Lightweight 1-50MB models
  • TFLite export - 3-10MB quantized vision models
  • ONNX export - Cross-platform deployment
  • GGUF quantization - Efficient LLM inference on CPU

๐Ÿ” Search & Embeddings

  • CLIP embeddings - Image-text similarity locally
  • RAG orchestration - Semantic image search
  • Vector DB ready - For semantic retrieval

๐Ÿ”ง Infrastructure

  • Prometheus metrics - Track latency, throughput
  • User auth - JWT-based with role-based access
  • No cloud required - Fully local, privacy-first
  • Docker support - docker-compose for quick setup

๐Ÿ—๏ธ Architecture

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚              CLIENT/USER LAYER                  โ”‚
โ”‚    Web UI / API / Mobile / Command Line         โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                   โ”‚
      โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
      โ”‚            โ”‚            โ”‚
โ”Œโ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”  โ”Œโ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”
โ”‚  USER   โ”‚  โ”‚  EDGE   โ”‚  โ”‚ SERVING  โ”‚
โ”‚  :8000  โ”‚  โ”‚  :8001  โ”‚  โ”‚ :8000    โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”˜  โ””โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”˜
      โ”‚ Auth       โ”‚ Training   โ”‚ Inference
      โ”‚            โ”‚  Inference โ”‚
      โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                   โ”‚
      โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
      โ”‚   JAX/Flax + XLA        โ”‚
      โ”‚  โ€ข JIT compilation      โ”‚
      โ”‚  โ€ข Graph optimization   โ”‚
      โ”‚  โ€ข 30%+ speedup         โ”‚
      โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                   โ”‚
      โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
      โ”‚  EXPORT PIPELINE        โ”‚
      โ”‚  TFLite | ONNX | SavedModel
      โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜
                   โ”‚
      โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
      โ”‚   DEPLOYMENT            โ”‚
      โ”‚  Mobile | Edge | Server  โ”‚
      โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

Key Layers:

  • User (8000): Signup, login, video generation, user endpoints
  • Edge (8001): Fine-tuning, detection, validation, inference
  • Serving (8000): FastAPI gateway, auth, metrics
  • Storage: SQLite model registry + checkpoint files

๐Ÿ“ Project Structure

open-ml-foundry/
โ”œโ”€โ”€ core/
โ”‚   โ””โ”€โ”€ session_store.py           # โœ… SQLite session + chat-event storage
โ”‚
โ”œโ”€โ”€ llm/                            # LLM Fine-Tuning
โ”‚   โ”œโ”€โ”€ supported_models.py        # โœ… Qwen, Gemma, DeepSeek, GLM, MiniMax, Kimi registry
โ”‚   โ”œโ”€โ”€ model_loader.py            # โœ… HF Hub + GGUF loading
โ”‚   โ”œโ”€โ”€ lora_trainer.py            # โœ… LoRA/QLoRA training (peft + transformers)
โ”‚   โ””โ”€โ”€ inference_engine.py        # โœ… Chat-style generation (HF + GGUF)
โ”‚
โ”œโ”€โ”€ edge/                           # Vision Fine-Tuning (Maintained)
โ”‚   โ”œโ”€โ”€ jax_train.py                # โœ… Training loop + CNN (run_finetuning)
โ”‚   โ”œโ”€โ”€ vision_session_adapter.py   # โœ… Wraps jax_train for session use
โ”‚   โ”œโ”€โ”€ model_registry.py           # โœ… Versioning + checkpoint storage
โ”‚   โ”œโ”€โ”€ vision_module.py            # โœ… FastAPI server (8001) โ€” legacy one-shot API
โ”‚   โ””โ”€โ”€ ...                         # Other vision services
โ”‚
โ”œโ”€โ”€ serving/
โ”‚   โ”œโ”€โ”€ main.py                     # โœ… FastAPI (8000), mounts session_api + /sessions UI
โ”‚   โ”œโ”€โ”€ session_api.py              # โœ… Unified session endpoints (LLM + vision)
โ”‚   โ”œโ”€โ”€ features_api.py             # โœ… Image gallery, datasets, training jobs, model registry, A/B testing
โ”‚   โ””โ”€โ”€ static/sessions_chat.html   # โœ… Chat-style session UI (vanilla JS, no build step)
โ”‚
โ”œโ”€โ”€ sessions.db                     # SQLite session store (gitignored, created on first run)
โ”œโ”€โ”€ requirements.txt                # โœ… Added: transformers, peft, accelerate, huggingface-hub
โ”œโ”€โ”€ docker-compose.yml              # โœ… Multi-service setup
โ”‚
โ”œโ”€โ”€ mlops/                          # ML Operations
โ”‚   โ”œโ”€โ”€ embedding_service.py       # โœ… CLIP embeddings
โ”‚   โ””โ”€โ”€ rag_orchestrator.py        # โœ… Semantic search
โ”‚
โ”œโ”€โ”€ docs/                           # Documentation
โ””โ”€โ”€ tests/                          # Unit tests

v0.3.0 Status:

  • โœ… Implemented: core/ (sessions), llm/ (LoRA training + inference wiring), edge/vision_session_adapter.py, serving/session_api.py, chat UI at /sessions
  • โณ Planned (v0.4.0): Multi-infrastructure support (Kubernetes, AWS, GCP, Azure, edge clusters)

See Known Limitations before deploying any of this.


๐Ÿ”Œ Core API Endpoints

Sessions (New โ€” unified LLM + vision)

  • GET /sessions - Chat-style session UI (web)
  • GET /api/v1/models - List supported LLM + vision models
  • POST /api/v1/sessions - Create session {name, model_type, model_name, model_format}
  • GET /api/v1/sessions - List sessions
  • GET /api/v1/sessions/{id} - Session details + full chat history
  • POST /api/v1/sessions/{id}/train - Start training (LoRA for LLM, run_finetuning for vision)
  • POST /api/v1/sessions/{id}/inference - Test the model, appended to history
  • POST /api/v1/sessions/{id}/note - Add a freeform note to the transcript
  • DELETE /api/v1/sessions/{id} - Delete session

Vision (Legacy, one-shot, no session)

  • POST /finetune - Train custom CNN on images
  • GET /detect - Real-time object detection
  • GET /download-model/{id} - Download trained checkpoint

Model Management:

  • GET /models-versions - View all model versions
  • GET /model-validation - Validation dashboard (web)
  • POST /validate-model/{model_id} - Run K-fold CV against a model's recorded training dataset

System:

  • GET /metrics - Prometheus metrics

No authentication layer โ€” this is a local, single-user tool (see Known Limitations).

Note: v0.3.0 is local-only. Multi-infrastructure support (Kubernetes, AWS, GCP, Azure) is planned for v0.4.0+.


๐Ÿ“ค Deployment

Local Server (Fastest)

docker-compose up
# Services on :8000 (API) and :8001 (Fine-tuning)
# 8-10ms inference latency

Mobile/Embedded (Offline)

# Export fine-tuned model
python -c "from edge.optimized_inference import OptimizedVisionInference; \
          engine = OptimizedVisionInference(); \
          engine.export_to_tflite('model.tflite'); \
          engine.export_to_onnx('model.onnx')"

# Result: 3-10MB quantized models for Android/iOS/Coral/RPi

Docker

docker build -t sentinel .
docker run -p 8000:8000 -p 8001:8001 sentinel

๐Ÿ“ˆ Performance

MetricValue
Inference latency8-10ms (GPU), 20-40ms (mobile)
XLA speedup30%+ throughput increase
Batch throughput200 images/sec (32-batch)
Model size (TFLite)3-10MB (quantized)
Training time5s-hours (depends on data)

โš™๏ธ Configuration

Environment Variables:

export JAX_PLATFORM_NAME=gpu      # Use GPU
export JAX_ENABLE_X64=False       # Use float32
export BATCH_SIZE=32              # Batch size
export DETECTION_THRESHOLD=0.7    # Object detection confidence

Dependencies:

  • JAX 0.4.13+, Flax 0.7+, PyTorch 2.0+, TensorFlow 2.12+, OpenCV 4.7+
  • Optional: pycoral (EdgeTPU), tf2onnx (ONNX), pyspark (Spark)

๐Ÿงช Testing & Benchmarking

# Run tests
pytest tests/test_security.py -v

# Benchmark XLA optimization
python edge/benchmark_xla.py

# Profile inference latency
python edge/optimized_inference.py

๐Ÿค Contributing

Contributions welcome! Areas needing work:

  • โŒ Model import UI (Streamlit/React)
  • โŒ Dataset browser & preview
  • โš ๏ธ Live training dashboard
  • โš ๏ธ Confusion matrix visualization
  • โœ… Inference optimization (help improve latency)
  • โœ… Mobile export testing

See CONTRIBUTING.md for guidelines.

๐ŸŽฏ Roadmap

v0.3.0 (Current) - Multi-Modal Foundations

  • โš ๏ธ LLM fine-tuning with LoRA/QLoRA (Qwen, Gemma, DeepSeek, GLM, MiniMax) โ€” see Known Limitations
  • โš ๏ธ Session-based training with chat history UI โ€” see Known Limitations
  • โš ๏ธ HuggingFace Hub + GGUF model support โ€” see Known Limitations
  • โœ… Vision models maintained (ResNet, YOLO, CLIP)
  • โœ… Chat UI shipped as a static page (serving/static/sessions_chat.html) โ€” plain HTML/JS, not React/Vue

v0.4.0 - Enhanced Sessions

  • Multi-turn dialogue training (dataset format already designed, needs real-world testing)
  • Inference optimization for local LLMs
  • Advanced hyperparameter tuning UI
  • Training progress streaming (WebSocket instead of 2s polling)
  • Model comparison dashboard

v0.5.0 - Multi-Modal Training

  • Unified JobSpec format (LLM + vision)
  • Multi-infrastructure support (Kubernetes, edge clusters)
  • Distributed training across devices
  • Cost tracking & optimization

v1.0.0 - Enterprise

  • Federated learning for privacy
  • Multi-cloud orchestration
  • Advanced audit logging & compliance
  • Custom backend plugins

No committed dates โ€” this is a volunteer-driven open-source project; see .reserve/ROADMAP.txt for effort estimates per phase.

๐Ÿ” Security

  • Authentication: JWT tokens with role-based access
  • Audit logging: Track all model training and inference
  • Input validation: Model & dataset shape checking
  • On-device inference: No data leaves your machine (optional)

See SECURITY.md for details.

โš ๏ธ Known Limitations

  • LLM sessions (core/, llm/, serving/session_api.py) have not been run end-to-end. The code has been reviewed but not executed against real dependencies โ€” treat it as "should work" until someone runs it and confirms.
  • No authentication layer. This is a local, single-user tool by design โ€” there's no login, no per-user access control, and no adversary model between "you" and "you." Don't expose it on a public network interface without adding your own access control in front of it.
  • LLM model repo ids are not all confirmed. See llm/supported_models.py โ€” each entry's verified field tracks whether its HuggingFace Hub repo id has actually been checked.
  • QLoRA and GGUF are optional installs. bitsandbytes (QLoRA, needs CUDA) and llama-cpp-python (GGUF) are commented out in requirements.txt because they need build tools or GPU hardware not present by default.

๐Ÿ“ License

MIT - see LICENSE

๐Ÿ™ Thanks To

JAX/Flax, XLA, PyTorch, TensorFlow, ONNX, and the open source ML community.


Fast local model fine-tuning with zero cloud dependency ๐Ÿš€