Open-source multi-modal fine-tuning. LLMs and vision models, fully local, accelerate edge deployments
All-in-one framework for fine-tuning large language models and vision models on your machine. Session-based training for Qwen3.8, GLM-5.3-Flash, Kimi K3, MiniMax-H3, DeepSeek-V4, Gemma 4 with LoRA/QLoRA. Plus vision models (ResNet, YOLO, CLIP). Privacy-first, runs completely local, deploy to edge devices. No cloud required.
๐ฏ Use Cases (What You Can Do Today)
LLM Fine-Tuning (Session-Based) - v0.3.0
| Use Case | Feature | Models |
|---|---|---|
| Fine-tune LLM | LoRA/QLoRA with session history | Qwen, Gemma, DeepSeek, GLM, MiniMax |
| Test inference | Chat-like interface during training | All LLM models |
| Model selection | Choose from HF Hub or GGUF models | HF Hub + GGUF |
| Multi-turn training | Train on dialogue/conversation data | All LLM models |
See Known Limitations โ this hasn't been run end-to-end yet.
Vision Fine-Tuning
| Use Case | Feature | Status |
|---|---|---|
| Custom image classifier | ResNet/CNN + your data (demo-scale training loop) | โ Working |
| Object detection | FasterRCNN ResNet50 (COCO) | โ Working |
| Mobile export | TFLite/ONNX + quantization | โ Working |
| Edge inference | 8-10ms latency | โ Working |
| Image search | CLIP embeddings | โ Working |
| Model versioning | Auto-versioning + rollback | โ Working |
| Test a fine-tuned checkpoint | Load checkpoint โ predict | โ
Working (FinetuneInference loads real params and runs a real forward pass) |
๐ Comparison with Existing Tools
LLM Fine-Tuning Focus
| Feature | Open ML Foundry | Unsloth | LLaMA-Factory | Axolotl |
|---|---|---|---|---|
| LoRA/QLoRA | โ | โ โญ (fastest) | โ | โ |
| Session-based UI | โ Chat history | โ ๏ธ Unsloth Studio (web UI + desktop app, has session features) | โ CLI/Web basic | โ CLI only |
| Model switching | โ New session per model (model_name is fixed at session creation, no code path changes it) | โ | โ ๏ธ Restart | โ ๏ธ Restart |
| Local-only (no cloud) | โ | โ | โ | โ |
| Popular models | Qwen, Gemma, DeepSeek, GLM | โญ Optimized | All HF models | All HF models |
| Inference UI | โ Chat interface | โ ๏ธ Unsloth Studio (chat + side-by-side model comparison) | โ ๏ธ Basic | โ |
| Vision + LLM | โ Both | โ LLM only | โ LLM only | โ LLM only |
Vision Fine-Tuning (Legacy Support)
| Feature | Open ML Foundry | PyTorch Lightning | FastAI |
|---|---|---|---|
| Image classification | โ ResNet/CNN | โ | โ |
| Object detection | โ FasterRCNN | โ | โ ๏ธ |
| Edge deployment | โ TFLite/ONNX | โ ๏ธ | โ |
| CLIP embeddings | โ | โ | โ |
TL;DR - Choose based on use case:
- Unsloth if: You want FASTEST LLM training speed (2-5x)
- Open ML Foundry if: You need LLM + vision + edge deployment + session-based UI
- LLaMA-Factory if: You want all HuggingFace models with advanced config
- PyTorch Lightning if: You need distributed/multi-GPU framework
๐ Documentation (Streamlined for Users)
Start experimenting with built-in models immediately โ no setup needed beyond pip install and docker-compose up.
Essential Guides
-
Models & Sample Experiments โ Start here
- Object detection on any image
- Train custom classifier in 5-10s
- Semantic image search
- 5 quick experiment ideas
-
CLI Guide โ Command-line interface (Phase 1)
sentinel model import- Import custom modelssentinel dataset prepare- Prepare training datasentinel train start- Train with built-in or custom models- Usage examples for all commands
-
- Install โ Start โ Use
-
- All endpoints
- Request/response examples
-
- How the system works
-
- Fine-tuning capabilities
- Model versioning
- Performance metrics
-
- XLA speedups
- Mobile export
- Benchmark results
-
- JWT tokens
- Role-based access
- Production hardening
๐ Get Started (5 Minutes)
1. Clone & Setup
git clone https://github.com/agentic-inquisit/open-ml-foundry.git
cd open-ml-foundry
# Run setup (checks dependencies, creates venv, installs packages)
./setup.sh
2. Session-Based Fine-Tuning (LLM + Vision, Chat UI)
# Start the API + session UI
uvicorn serving.main:app --reload --port 8000
# Open http://localhost:8000/sessions in a browser:
# 1. "+ New Session" โ pick LLM (Qwen, Gemma, DeepSeek, GLM, MiniMax) or Vision
# 2. Point at a dataset (JSONL for LLM, image path for vision)
# 3. Start Training โ progress streams into the session as chat events
# 4. Test tab โ send a prompt/image, see the model's reply in the same thread
# 5. Every session keeps its full history โ switch sessions from the sidebar
Or drive the same flow over the REST API directly โ see Core API Endpoints below.
3. Vision Fine-Tuning (Direct API, no session)
# Object detection (pretrained, no fine-tuning needed)
curl http://localhost:8001/detect -F "image=@test.jpg"
# One-shot fine-tune (legacy endpoint, still works outside sessions)
curl -X POST http://localhost:8001/finetune -F "dataset=@image.jpg" -F "target_object=my_class"
4. Full Server Setup
# Terminal 1: Start all services
./start.sh
# Terminal 2: Test LLM inference
curl -X POST http://localhost:8000/llm/inference \
-H "Content-Type: application/json" \
-d '{"session_id": "my-session", "prompt": "Hello, how are you?"}'
# Test vision inference
curl http://localhost:8001/detect -F "image=@test.jpg"
โก Core Features
๐ฌ LLM Fine-Tuning (Session-Based) - v0.3.0
Implemented in core/session_store.py, llm/, serving/session_api.py, serving/static/sessions_chat.html. Not yet run end-to-end โ see Known Limitations.
- Session management - Each fine-tuning is a session with persistent, chat-like history (SQLite)
- LoRA/QLoRA - via
peft+transformers; QLoRA needsbitsandbytes+ CUDA GPU (not installed by default) - Model support - Qwen, Gemma, DeepSeek, GLM, MiniMax, Kimi (via HF Hub โ see
llm/supported_models.pyfor repo id verification status) - Model formats - HuggingFace Hub + GGUF (GGUF needs
llama-cpp-python, not installed by default โ build-tool dependency) - Inference in-session - Test the model mid-conversation from the same chat UI
- Multi-turn support - JSONL datasets with
messages(chat) orprompt/completionshape - Checkpointing - LoRA adapter saved per session under
training_outputs/llm/<session_id>/
๐ผ๏ธ Vision Fine-Tuning (Maintained)
- Image classification - Fine-tune ResNet/CNN on your images
- Object detection - FasterRCNN ResNet50+FPN (COCO pretrained)
- Quick training - 5-10s for demo datasets, scales to hours
- Validation & early stopping - Automatic train/val split with patience
- Auto-versioning - v1.0, v1.1, v2.0 format with rollback
๐ Model Management
- Performance tracking - Loss, accuracy, metrics per epoch in SQLite
- Model comparison - Compare versions side-by-side
- Session history - View all training runs like conversations
- Checkpoint storage - Save/restore at any epoch
๐ Optimization & Export
- LoRA adapter export - Lightweight 1-50MB models
- TFLite export - 3-10MB quantized vision models
- ONNX export - Cross-platform deployment
- GGUF quantization - Efficient LLM inference on CPU
๐ Search & Embeddings
- CLIP embeddings - Image-text similarity locally
- RAG orchestration - Semantic image search
- Vector DB ready - For semantic retrieval
๐ง Infrastructure
- Prometheus metrics - Track latency, throughput
- User auth - JWT-based with role-based access
- No cloud required - Fully local, privacy-first
- Docker support - docker-compose for quick setup
๐๏ธ Architecture
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ CLIENT/USER LAYER โ
โ Web UI / API / Mobile / Command Line โ
โโโโโโโโโโโโโโโโโโโโฌโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
โโโโโโโโโโโโโโผโโโโโโโโโโโโโ
โ โ โ
โโโโโโโผโโโโ โโโโโโโผโโโโ โโโโโโผโโโโโ
โ USER โ โ EDGE โ โ SERVING โ
โ :8000 โ โ :8001 โ โ :8000 โ
โโโโโโโฌโโโโ โโโโโโโฌโโโโ โโโโโโฌโโโโโโ
โ Auth โ Training โ Inference
โ โ Inference โ
โโโโโโโโโโโโโโผโโโโโโโโโโโโโ
โ
โโโโโโโโโโโโโโผโโโโโโโโโโโโโ
โ JAX/Flax + XLA โ
โ โข JIT compilation โ
โ โข Graph optimization โ
โ โข 30%+ speedup โ
โโโโโโโโโโโโโโฌโโโโโโโโโโโโโ
โ
โโโโโโโโโโโโโโผโโโโโโโโโโโโโ
โ EXPORT PIPELINE โ
โ TFLite | ONNX | SavedModel
โโโโโโโโโโโโโโฌโโโโโโโโโโโโโ
โ
โโโโโโโโโโโโโโผโโโโโโโโโโโโโ
โ DEPLOYMENT โ
โ Mobile | Edge | Server โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโ
Key Layers:
- User (8000): Signup, login, video generation, user endpoints
- Edge (8001): Fine-tuning, detection, validation, inference
- Serving (8000): FastAPI gateway, auth, metrics
- Storage: SQLite model registry + checkpoint files
๐ Project Structure
open-ml-foundry/
โโโ core/
โ โโโ session_store.py # โ
SQLite session + chat-event storage
โ
โโโ llm/ # LLM Fine-Tuning
โ โโโ supported_models.py # โ
Qwen, Gemma, DeepSeek, GLM, MiniMax, Kimi registry
โ โโโ model_loader.py # โ
HF Hub + GGUF loading
โ โโโ lora_trainer.py # โ
LoRA/QLoRA training (peft + transformers)
โ โโโ inference_engine.py # โ
Chat-style generation (HF + GGUF)
โ
โโโ edge/ # Vision Fine-Tuning (Maintained)
โ โโโ jax_train.py # โ
Training loop + CNN (run_finetuning)
โ โโโ vision_session_adapter.py # โ
Wraps jax_train for session use
โ โโโ model_registry.py # โ
Versioning + checkpoint storage
โ โโโ vision_module.py # โ
FastAPI server (8001) โ legacy one-shot API
โ โโโ ... # Other vision services
โ
โโโ serving/
โ โโโ main.py # โ
FastAPI (8000), mounts session_api + /sessions UI
โ โโโ session_api.py # โ
Unified session endpoints (LLM + vision)
โ โโโ features_api.py # โ
Image gallery, datasets, training jobs, model registry, A/B testing
โ โโโ static/sessions_chat.html # โ
Chat-style session UI (vanilla JS, no build step)
โ
โโโ sessions.db # SQLite session store (gitignored, created on first run)
โโโ requirements.txt # โ
Added: transformers, peft, accelerate, huggingface-hub
โโโ docker-compose.yml # โ
Multi-service setup
โ
โโโ mlops/ # ML Operations
โ โโโ embedding_service.py # โ
CLIP embeddings
โ โโโ rag_orchestrator.py # โ
Semantic search
โ
โโโ docs/ # Documentation
โโโ tests/ # Unit tests
v0.3.0 Status:
- โ
Implemented: core/ (sessions), llm/ (LoRA training + inference wiring), edge/vision_session_adapter.py, serving/session_api.py, chat UI at
/sessions - โณ Planned (v0.4.0): Multi-infrastructure support (Kubernetes, AWS, GCP, Azure, edge clusters)
See Known Limitations before deploying any of this.
๐ Core API Endpoints
Sessions (New โ unified LLM + vision)
GET /sessions- Chat-style session UI (web)GET /api/v1/models- List supported LLM + vision modelsPOST /api/v1/sessions- Create session{name, model_type, model_name, model_format}GET /api/v1/sessions- List sessionsGET /api/v1/sessions/{id}- Session details + full chat historyPOST /api/v1/sessions/{id}/train- Start training (LoRA for LLM,run_finetuningfor vision)POST /api/v1/sessions/{id}/inference- Test the model, appended to historyPOST /api/v1/sessions/{id}/note- Add a freeform note to the transcriptDELETE /api/v1/sessions/{id}- Delete session
Vision (Legacy, one-shot, no session)
POST /finetune- Train custom CNN on imagesGET /detect- Real-time object detectionGET /download-model/{id}- Download trained checkpoint
Model Management:
GET /models-versions- View all model versionsGET /model-validation- Validation dashboard (web)POST /validate-model/{model_id}- Run K-fold CV against a model's recorded training dataset
System:
GET /metrics- Prometheus metrics
No authentication layer โ this is a local, single-user tool (see Known Limitations).
Note: v0.3.0 is local-only. Multi-infrastructure support (Kubernetes, AWS, GCP, Azure) is planned for v0.4.0+.
๐ค Deployment
Local Server (Fastest)
docker-compose up
# Services on :8000 (API) and :8001 (Fine-tuning)
# 8-10ms inference latency
Mobile/Embedded (Offline)
# Export fine-tuned model
python -c "from edge.optimized_inference import OptimizedVisionInference; \
engine = OptimizedVisionInference(); \
engine.export_to_tflite('model.tflite'); \
engine.export_to_onnx('model.onnx')"
# Result: 3-10MB quantized models for Android/iOS/Coral/RPi
Docker
docker build -t sentinel .
docker run -p 8000:8000 -p 8001:8001 sentinel
๐ Performance
| Metric | Value |
|---|---|
| Inference latency | 8-10ms (GPU), 20-40ms (mobile) |
| XLA speedup | 30%+ throughput increase |
| Batch throughput | 200 images/sec (32-batch) |
| Model size (TFLite) | 3-10MB (quantized) |
| Training time | 5s-hours (depends on data) |
โ๏ธ Configuration
Environment Variables:
export JAX_PLATFORM_NAME=gpu # Use GPU
export JAX_ENABLE_X64=False # Use float32
export BATCH_SIZE=32 # Batch size
export DETECTION_THRESHOLD=0.7 # Object detection confidence
Dependencies:
- JAX 0.4.13+, Flax 0.7+, PyTorch 2.0+, TensorFlow 2.12+, OpenCV 4.7+
- Optional: pycoral (EdgeTPU), tf2onnx (ONNX), pyspark (Spark)
๐งช Testing & Benchmarking
# Run tests
pytest tests/test_security.py -v
# Benchmark XLA optimization
python edge/benchmark_xla.py
# Profile inference latency
python edge/optimized_inference.py
๐ค Contributing
Contributions welcome! Areas needing work:
- โ Model import UI (Streamlit/React)
- โ Dataset browser & preview
- โ ๏ธ Live training dashboard
- โ ๏ธ Confusion matrix visualization
- โ Inference optimization (help improve latency)
- โ Mobile export testing
See CONTRIBUTING.md for guidelines.
๐ฏ Roadmap
v0.3.0 (Current) - Multi-Modal Foundations
- โ ๏ธ LLM fine-tuning with LoRA/QLoRA (Qwen, Gemma, DeepSeek, GLM, MiniMax) โ see Known Limitations
- โ ๏ธ Session-based training with chat history UI โ see Known Limitations
- โ ๏ธ HuggingFace Hub + GGUF model support โ see Known Limitations
- โ Vision models maintained (ResNet, YOLO, CLIP)
- โ
Chat UI shipped as a static page (
serving/static/sessions_chat.html) โ plain HTML/JS, not React/Vue
v0.4.0 - Enhanced Sessions
- Multi-turn dialogue training (dataset format already designed, needs real-world testing)
- Inference optimization for local LLMs
- Advanced hyperparameter tuning UI
- Training progress streaming (WebSocket instead of 2s polling)
- Model comparison dashboard
v0.5.0 - Multi-Modal Training
- Unified JobSpec format (LLM + vision)
- Multi-infrastructure support (Kubernetes, edge clusters)
- Distributed training across devices
- Cost tracking & optimization
v1.0.0 - Enterprise
- Federated learning for privacy
- Multi-cloud orchestration
- Advanced audit logging & compliance
- Custom backend plugins
No committed dates โ this is a volunteer-driven open-source project; see .reserve/ROADMAP.txt for effort estimates per phase.
๐ Security
- Authentication: JWT tokens with role-based access
- Audit logging: Track all model training and inference
- Input validation: Model & dataset shape checking
- On-device inference: No data leaves your machine (optional)
See SECURITY.md for details.
โ ๏ธ Known Limitations
- LLM sessions (
core/,llm/,serving/session_api.py) have not been run end-to-end. The code has been reviewed but not executed against real dependencies โ treat it as "should work" until someone runs it and confirms. - No authentication layer. This is a local, single-user tool by design โ there's no login, no per-user access control, and no adversary model between "you" and "you." Don't expose it on a public network interface without adding your own access control in front of it.
- LLM model repo ids are not all confirmed. See
llm/supported_models.pyโ each entry'sverifiedfield tracks whether its HuggingFace Hub repo id has actually been checked. - QLoRA and GGUF are optional installs.
bitsandbytes(QLoRA, needs CUDA) andllama-cpp-python(GGUF) are commented out inrequirements.txtbecause they need build tools or GPU hardware not present by default.
๐ License
MIT - see LICENSE
๐ Thanks To
JAX/Flax, XLA, PyTorch, TensorFlow, ONNX, and the open source ML community.
Fast local model fine-tuning with zero cloud dependency ๐