Machina: open machine intelligence you can audit
Why Machina
TL;DR. Machina is an open harness for machine intelligence: it turns a window of sensor data into a diagnosis, a remaining-life estimate, an energy or process-quality finding and an evidence brief, and it does so with models you can read, abstention you can tune and an audit trail you can trust. The ethos fits one line: signals stay legible, models stay portable, actions stay governed.
Industrial machine learning usually fails for dull reasons: a model that is accurate on a lab bench and unusable on a factory floor, a score nobody can explain to the engineer who has to act on it, a log that stores too much, a pipeline that can silently skip a human. Machina is a bet that most of the value is in the plumbing around the model: honest abstention, inspectable features, replaceable parts, and a boundary where the software stops and a qualified person starts.
MACHINA_DB_PATH is set). Source: github.com/Clerktree/machina-intelligence · src/machina_harness/api.py, mcp_server.py, platform.py| Principle | What it means in the code |
|---|---|
| Signals stay legible | Hand-readable features (RMS, kurtosis, crest factor, band energies, envelope frequency) listed in metadata and mirrored between training and inference; an explainable robust-score anomaly rule that names the sensors that drove it; metadata-only audit rows that never store raw sensor values. |
| Models stay portable | Plain scikit-learn artifacts with a model card each; a registry of replaceable “capability plugins”; path overrides via environment variables; SHA-256 of every artifact on /v1/model-health; CPU-only inference with no network calls; a small LoRA adapter for the agent instead of a bespoke foundation model. |
| Actions stay governed | Decision support, never a control loop; human_review_required is hard-coded to true on every finding; the classifier abstains below 0.65 confidence; API-key middleware and request ids; a hardened container; a written readiness gate that refuses to claim certification. |
Bearing faults: from 0.59 to 0.99, and why to distrust 0.99
The first capability classifies a vibration window as normal, ball, inner-race or outer-race fault, using the Case Western Reserve University bearing dataset (the raw files are kept outside the repository). Two checkpoints ship: a four-feature RandomForest baseline, and the enhanced model used by the container.
| Class | Baseline P / R / F1 | Shipped P / R / F1 | Shipped support |
|---|---|---|---|
| ball | 0.504 / 0.725 / 0.595 | 1.000 / 0.938 / 0.968 | 48 |
| inner race | 0.627 / 0.650 / 0.638 | 1.000 / 1.000 / 1.000 | 80 |
| normal | 1.000 / 0.250 / 0.400 | 1.000 / 1.000 / 1.000 | 16 |
| outer race | 0.805 / 0.644 / 0.715 | 0.974 / 1.000 / 0.987 | 112 |
| accuracy · macro-F1 | 0.655 · 0.587 | 0.988 · 0.989 | 256 windows |
Does it survive a speed it has not seen?
A random split lets nearly identical windows sit on both sides. The enhanced run therefore splits by file (whole recordings held out) and also by motor speed: leave-one-RPM-out over the dataset’s four speeds (1730, 1750, 1772, 1797 rpm).
Anomaly detection and energy: rules you can check by hand
Not everything should be a model. The anomaly detector is a robust z-score of the latest value against the window’s median and median absolute deviation, mapped to a 0–1 score per sensor; the window score is the worst sensor, and the three most responsible sensors are reported. Energy analytics is the same idea for efficiency: energy per unit of output relative to a median baseline.
score = clip((z − 1.5) / 4.5, 0, 1). The window score is the maximum over sensors; at least 0.75 is “critical”, at least 0.35 “watch”. The two z cut-offs shown (about 3.1 and 4.9) are derived from the thresholds. No learning is involved, which is the point: an engineer can check it by hand. Source: github.com/Clerktree/machina-intelligence · src/machina_harness/anomaly.py| Rule | |
|---|---|
| Baseline | median of power ÷ output rate over at least 3 history samples |
| Ratio | current ÷ baseline |
| Score | clip((ratio − 1.05) / 0.95, 0, 1) |
| Status | critical ≥ 0.55 · watch ≥ 0.15 · otherwise normal |
| Example (from the tests) | 100 kW at 10 units history; 160 kW now → critical; 102 kW → normal |
Remaining useful life: the honest number is 62 cycles
The second capability estimates how many cycles an engine has left, trained on NASA’s C-MAPSS turbofan degradation simulations (subset FD001). At runtime it returns a point estimate with an interval taken from the spread of the forest’s individual trees.
| Item | Value |
|---|---|
| Dataset | NASA C-MAPSS FD001: official train, test and RUL files |
| Features (17) | cycle, cycle fraction, and the 15 non-constant sensors (2, 3, 4, 6, 7, 8, 9, 11, 12, 13, 14, 15, 17, 20, 21); operating settings not used |
| Model | ExtraTreesRegressor · 250 trees · min leaf 2 · max features 0.8 |
| Target | RUL = engine max cycle − cycle, clipped at 125 |
| Evaluation | last cycle of each of 100 test engines against the official RUL file (not clipped) |
| Uncertainty at runtime | 1.96 × the standard deviation of per-tree predictions, at least 1 cycle: a transparent proxy for a prediction interval |
| Artifact size | 161 MB (the largest of the four checkpoints) |
cycle_fraction feature is cycle divided by the engine’s total life during training, which leaks the answer, while at test time it is computed from the truncated observed trajectory. Source: github.com/Clerktree/machina-intelligence · artifacts/rul-cmapss/metadata.json; docs/RUL_BASELINE.mdProcess quality: accuracy 98%, macro-F1 55%
The third capability predicts a process failure mode from six machine settings, using the synthetic UCI AI4I 2020 dataset: tool wear, heat dissipation, power, overstrain or random failure, as a probability that something other than “normal” is happening.
The agent: routing, not reasoning from scratch
Machina’s design splits intelligence in two. Specialist models operate on machine signals; a small language model translates a natural-language request into the right tool call. The agent is a Mistral-7B-Instruct adapter, trained with QLoRA to emit one structured tool call.
| Setting | Value |
|---|---|
| Base model | mistralai/Mistral-7B-Instruct-v0.3 (Apache-2.0, native function calling) |
| Method | QLoRA: 4-bit NF4, double quantisation, bf16 compute |
| LoRA | rank 32 · alpha 64 · dropout 0.05 · q, k, v, o, gate, up, down projections |
| Optimisation | learning rate 2e-4, cosine, warm-up 5% · batch 1 × accumulation 8 · 2 epochs · paged AdamW 8-bit |
| Sequence length | 1,280 tokens, gradient checkpointing |
| Loss | only on the first assistant tool-call turn: the prompt and tool schema are masked |
| Data | 1,200 synthetic examples: 10 tools × 120; 92 / 8 train / eval split |
| Hardware | one lab GPU (RTX 4500 Ada) |
The training set is a synthetic, grounded routing set: it teaches routing and response style, not machine facts. It contains 1,200 examples, exactly 120 for each of ten tools, built from 50 distinct user strings. That makes it a good fit for a router and a poor basis for claims about generalisation, which is why the repository says to validate the adapter on held-out routing examples before calling it production-ready.
The 15 MCP tools, and the 10 the router is trained on
| MCP tool | In the router’s training schema |
|---|---|
machine_harness_health | yes |
list_machine_intelligence_capabilities | yes |
machina_platform_snapshot | yes |
list_machine_assets | yes |
list_registered_models | yes |
search_machine_knowledge | yes |
prepare_maintenance_brief | yes |
estimate_remaining_useful_life | yes |
analyze_machine_energy | yes |
predict_process_quality | yes |
analyze_machine_window | no |
register_machine_asset | no |
record_machine_telemetry | no |
record_maintenance_event | no |
index_maintenance_document | no |
analyze_machine_window) and the write tools are not in the trained schema yet. Source: src/machina_harness/mcp_server.py, scripts/build_agent_dataset.pyEvidence for the language layer
The “reasoning” side of the platform is a retrieval brief. Maintenance documents are indexed in an SQLite full-text table and searched with BM25; prepare_maintenance_brief returns the asset record, the top evidence passages, the registered model ids and three fixed instructions for whatever language model consumes them:
Use only the returned evidence and model outputs as factual support.
Separate observed signals from hypotheses.
Recommend qualified human inspection for safety-critical decisions.Governance as code
The least glamorous part of the repository is the part most worth copying. The platform treats every analysis as something that may be audited, challenged and overridden.
X-Request-ID is echoed or generated on every response and logged with latency. Source: github.com/Clerktree/machina-intelligence · src/machina_harness/api.py, platform.py| Control | How |
|---|---|
| Authentication | optional MACHINA_API_KEY; when set, every route except /health needs X-Machina-API-Key (401 otherwise) |
| Deployment | the compose file refuses to start without the key, runs read-only with a tmpfs /tmp, no-new-privileges, all capabilities dropped, health-checked every 30 s |
| Abstention | fault confidence below MACHINA_FAULT_MIN_CONFIDENCE (default 0.65) returns no diagnosis |
| Human in the loop | human_review_required = True on every finding; critical findings advise stopping reliance on the automated diagnosis |
| Traceability | SHA-256 of every model artifact on /v1/model-health; /ready reports “degraded” if a plugin is missing |
| Release gate | verify_release.py requires README, metadata and weights per artifact and, for the enhanced bearing model, a leave-one-speed-out floor and a “not evidence of generalization” sentence in its model card |
Portability and the edge
“Portable” here has a concrete meaning: the models are plain scikit-learn joblib files with a metadata JSON and a model card, they are registered as plugins that can be swapped, and path overrides let a deployment point to its own checkpoints. They run on a CPU with no network access at inference.
python:3.12-slim. Source: github.com/Clerktree/machina-intelligence · artifacts/*/model.joblib (Git LFS pointer sizes), DockerfileRun it
Python 3.10 or newer. Fetch the large files with Git LFS, install the package and start the API; the interactive documentation is at /docs.
git clone https://github.com/Clerktree/machina-intelligence && cd machina-intelligence
git lfs pull # model.joblib files are Git LFS objects
python3 -m venv .venv && source .venv/bin/activate
pip install -e '.[dev]' # extras: dev, train, mcp, runtime, agent
uvicorn machina_harness.api:app --reload # http://127.0.0.1:8000/docs
curl -X POST http://127.0.0.1:8000/v1/analyze \
-H 'content-type: application/json' -d @configs/sample-window.json
curl http://127.0.0.1:8000/v1/model-health
# container (the key is mandatory)
export MACHINA_API_KEY="replace-with-a-long-random-secret"
docker compose up --build -d
curl http://127.0.0.1:8000/ready -H "X-Machina-API-Key: $MACHINA_API_KEY"
# MCP over stdio (e.g. for Grok Build)
pip install -e '.[mcp]' && python3 -m machina_harness.mcp_serverA few practical notes. The classifier imports scipy.signal.hilbert, which arrives through scikit-learn rather than being declared on its own. The model cards declare Apache-2.0. The capability registry is the source of truth for what is live: fault diagnosis, anomaly detection, remaining useful life, energy intelligence and quality prediction are available, while the maintenance copilot and machine knowledge capabilities are marked as planned.
The decision-support framing is repeated in the repository, and it is worth repeating here: the model is decision support, not a safety controller. Human inspection and site-specific validation are required before any maintenance action.


