Machina: open machine intelligence you can audit

Listen
Share

Why Machina

TL;DR. Machina is an open harness for machine intelligence: it turns a window of sensor data into a diagnosis, a remaining-life estimate, an energy or process-quality finding and an evidence brief, and it does so with models you can read, abstention you can tune and an audit trail you can trust. The ethos fits one line: signals stay legible, models stay portable, actions stay governed.

15MCP tools over one inference function
21hand-readable features in the shipped bearing model
0.65confidence below which the classifier abstains
0raw sensor values ever stored in the audit log

Industrial machine learning usually fails for dull reasons: a model that is accurate on a lab bench and unusable on a factory floor, a score nobody can explain to the engineer who has to act on it, a log that stores too much, a pipeline that can silently skip a human. Machina is a bet that most of the value is in the plumbing around the model: honest abstention, inspectable features, replaceable parts, and a boundary where the software stops and a qualified person starts.

Machina architectureClients (HTTP, MCP, agent) reach one FastAPI platform with an inference core of five capabilities; state lives in an optional SQLite database.HTTP clientPOST /v1/analyzeMCP clientstdio · 15 toolsMachina agentMistral-7B + LoRA routertool callsmachina_harness · FastAPIAPI key · request idX-Machina-API-Key · X-Request-IDone inference functionthe same code path for HTTP and MCPfault diagnosisCWRU · calibrated ETanomaly detectionrobust score · rulesremaining useful lifeC-MAPSS · ExtraTreesenergy intelligencebaseline ratio · rulesprocess qualityAI4I · ExtraTreesmaintenance briefFTS5 / BM25 evidencemodels: scikit-learn joblib + metadata.json per artifactSQLite · MACHINA_DB_PATHassetsregistered machinestelemetryper assetmaintenance_eventswork historymodelsplugin registryknowledge_ftsFTS5 · BM25 searchinference_auditsmetadata only
Figure 1 The whole platform in one picture. Everything funnels through a single inference function, so the HTTP API and the MCP tools cannot silently use different model paths; state is optional (in memory by default, SQLite when MACHINA_DB_PATH is set). Source: github.com/Clerktree/machina-intelligence · src/machina_harness/api.py, mcp_server.py, platform.py
PrincipleWhat it means in the code
Signals stay legibleHand-readable features (RMS, kurtosis, crest factor, band energies, envelope frequency) listed in metadata and mirrored between training and inference; an explainable robust-score anomaly rule that names the sensors that drove it; metadata-only audit rows that never store raw sensor values.
Models stay portablePlain scikit-learn artifacts with a model card each; a registry of replaceable “capability plugins”; path overrides via environment variables; SHA-256 of every artifact on /v1/model-health; CPU-only inference with no network calls; a small LoRA adapter for the agent instead of a bespoke foundation model.
Actions stay governedDecision support, never a control loop; human_review_required is hard-coded to true on every finding; the classifier abstains below 0.65 confidence; API-key middleware and request ids; a hardened container; a written readiness gate that refuses to claim certification.
The three design commitments, and where each one lives. The mapping is a reading of the code against the stated principles. Source: README.md, docs/MACHINA_PLATFORM.md, docs/INDUSTRY_READINESS.md
Platform layersFive layers: perception, machine state, prediction, reasoning and action.PerceptionsensorssignalsMachine stateassetstelemetryPredictionfaults · RULenergy · qualityReasoningevidencetool routingActionalerts · reportsfeedbackAction never silently controls safety-critical equipment: it drafts alerts, work orders and reports for a human.
Figure 2 The five layers of the platform design. The first release builds the Prediction layer (three models plus two rule-based analyses) and the plumbing of the others. Source: github.com/Clerktree/machina-intelligence · docs/MACHINA_PLATFORM.md

Bearing faults: from 0.59 to 0.99, and why to distrust 0.99

The first capability classifies a vibration window as normal, ball, inner-race or outer-race fault, using the Case Western Reserve University bearing dataset (the raw files are kept outside the repository). Two checkpoints ship: a four-feature RandomForest baseline, and the enhanced model used by the container.

Bearing-fault classification pipelineA vibration window is conditioned, turned into 21 features, classified by a calibrated ExtraTrees model, and passed through a confidence gate that outputs a class or abstains.Vibrationwindow4,096 samples12 kHz · 0.34 sConditionmean-centreHann window21 featurestime · spectrumenvelope · bandsCalibratedExtraTrees600 treessigmoid · cv = 3Confidencegatep_max ≥ 0.65env-tunablepredicted classball · inner raceouter race · normalabstainhuman reviewrequired≥ 0.65< 0.65Signals stay legible: every feature is an inspectable quantity.Models stay portable: a joblib file and a JSON of its features.Actions stay governed: low confidence is an answer, not a guess.
Figure 3 From a vibration window to a diagnosis, or an honest “I am not sure”. The runtime classifier returns the most probable class only when its probability reaches 0.65 (tunable by environment variable); otherwise it abstains and recommends collecting a longer, better signal and requiring qualified review. Source: github.com/Clerktree/machina-intelligence · src/machina_harness/classifier.py
The 21 features of the shipped bearing model, by familyThe 21 features of the shipped bearing model, by familytime domainrms · peak · kurtosis · crest · std …10spectral shapedominant frequency · centroid · entropy3envelopeHilbert-envelope dominant freq.1band energy7 power fractions up to Nyquist7
Figure 4 Every feature is a quantity an engineer can reason about. The envelope-spectrum frequency is the classic bearing-diagnosis trick: demodulate the vibration so the fault’s repetition rate shows up as a line. Source: github.com/Clerktree/machina-intelligence · scripts/train_cwru_enhanced.py, src/machina_harness/classifier.py
Macro-F1 on a grouped test split (whole files held out)Macro-F1 on a grouped test split (whole files held out)Baseline RF · 4 features161 files · 1,288 windows · 328 test0.587ExtraTrees · 21 features64 files · 1,024 windows · 256 test0.996RandomForest · 21 featuressame split0.997Calibrated ExtraTrees (shipped)same split · sigmoid, cv = 30.989
Figure 5 What better features bought. The 4-feature baseline reaches macro-F1 0.587; 21 signal-processing features lift every 21-feature model to about 0.99. The two datasets and splits differ (the enhanced run uses 12 kHz drive-end files only), so read the jump as a direction, not a controlled ablation. Source: github.com/Clerktree/machina-intelligence · artifacts/cwru-baseline and cwru-enhanced metadata.json
ClassBaseline P / R / F1Shipped P / R / F1Shipped support
ball0.504 / 0.725 / 0.5951.000 / 0.938 / 0.96848
inner race0.627 / 0.650 / 0.6381.000 / 1.000 / 1.00080
normal1.000 / 0.250 / 0.4001.000 / 1.000 / 1.00016
outer race0.805 / 0.644 / 0.7150.974 / 1.000 / 0.987112
accuracy · macro-F10.655 · 0.5870.988 · 0.989256 windows
Per-class results. The baseline’s weakness is plain in the table: it misses three quarters of the (only eight) normal windows. The repository’s own verdict on that baseline: “The normal recall is inadequate for production.” Source: artifacts/cwru-baseline/metadata.json, artifacts/cwru-enhanced/metadata.json
Baseline confusion (true rows, predicted columns)Baseline confusion (true rows, predicted columns)ballinnernormalouterball581507inner race1052018normal6020outer race41160103
Shipped model confusion (true rows, predicted columns)Shipped model confusion (true rows, predicted columns)ballinnernormalouterball45003inner race08000normal00160outer race000112
Figure 6 Confusion matrices, baseline (top) and shipped calibrated model (bottom), in windows. The baseline smears outer-race faults into ball faults (41 of 160); the shipped model’s only errors are three ball windows read as outer race. Source: github.com/Clerktree/machina-intelligence · artifacts/*/metadata.json

Does it survive a speed it has not seen?

A random split lets nearly identical windows sit on both sides. The enhanced run therefore splits by file (whole recordings held out) and also by motor speed: leave-one-RPM-out over the dataset’s four speeds (1730, 1750, 1772, 1797 rpm).

Leave-one-RPM-out macro-F1 by held-out speedLeave-one-RPM-out macro-F1 by held-out speed0.9800.9850.9900.9951.0001730175017721797ExtraTreesRandomForestCalibrated ET (shipped)held-out operating speed (rpm); the model never saw that speed during training
Figure 7 A harder test than a random split: train on three motor speeds, test on the fourth. The shipped model holds macro-F1 between 0.984 and 0.997 (mean 0.9915 on the repository’s own accounting). The selection rule is explicit in the training script: the calibrated model is chosen when its mean leave-one-speed-out score is at least 0.98 and its minimum is at least 0.95, even though it is not the top scorer on the grouped split, because calibrated probabilities are what the abstention gate needs. Source: github.com/Clerktree/machina-intelligence · artifacts/cwru-enhanced/metadata.json; scripts/train_cwru_enhanced.py

Anomaly detection and energy: rules you can check by hand

Not everything should be a model. The anomaly detector is a robust z-score of the latest value against the window’s median and median absolute deviation, mapped to a 0–1 score per sensor; the window score is the worst sensor, and the three most responsible sensors are reported. Energy analytics is the same idea for efficiency: energy per unit of output relative to a median baseline.

Anomaly score as a function of robust zA piecewise-linear score that is zero until the robust z-score reaches 1.5 and one from 6.0; watch above 0.35 and critical above 0.75.watch ≥ 0.35critical ≥ 0.75z ≈ 3.1z ≈ 4.900.5101.534.568robust z of the latest value: |last − median| / (1.4826 · MAD)score
Figure 8 The anomaly rule: a robust z-score of the most recent value against the window’s median and MAD, mapped by score = clip((z − 1.5) / 4.5, 0, 1). The window score is the maximum over sensors; at least 0.75 is “critical”, at least 0.35 “watch”. The two z cut-offs shown (about 3.1 and 4.9) are derived from the thresholds. No learning is involved, which is the point: an engineer can check it by hand. Source: github.com/Clerktree/machina-intelligence · src/machina_harness/anomaly.py
Rule
Baselinemedian of power ÷ output rate over at least 3 history samples
Ratiocurrent ÷ baseline
Scoreclip((ratio − 1.05) / 0.95, 0, 1)
Statuscritical ≥ 0.55 · watch ≥ 0.15 · otherwise normal
Example (from the tests)100 kW at 10 units history; 160 kW now → critical; 102 kW → normal
Energy analytics is a transparent rule, not a learned model: energy per unit of output relative to a robust baseline. Source: src/machina_harness/energy.py, tests/test_energy.py

Remaining useful life: the honest number is 62 cycles

The second capability estimates how many cycles an engine has left, trained on NASA’s C-MAPSS turbofan degradation simulations (subset FD001). At runtime it returns a point estimate with an interval taken from the spread of the forest’s individual trees.

The clipped remaining-useful-life targetA degradation schematic: the true remaining life falls linearly to zero, while the training target is capped at 125 cycles for the early part of an engine’s life.target = min(cycles left, 125)cycles left (the two lines coincide below 125)0125200engine life (cycles) → failureRUL (cycles)
Figure 9 Why the RUL target is capped. Early in an engine’s life nothing measurable distinguishes 190 cycles left from 130, so the common C-MAPSS convention caps the target at 125 and learns the interesting end of the curve. This is a schematic of that convention, not engine data. Source: github.com/Clerktree/machina-intelligence · scripts/train_rul.py
ItemValue
DatasetNASA C-MAPSS FD001: official train, test and RUL files
Features (17)cycle, cycle fraction, and the 15 non-constant sensors (2, 3, 4, 6, 7, 8, 9, 11, 12, 13, 14, 15, 17, 20, 21); operating settings not used
ModelExtraTreesRegressor · 250 trees · min leaf 2 · max features 0.8
TargetRUL = engine max cycle − cycle, clipped at 125
Evaluationlast cycle of each of 100 test engines against the official RUL file (not clipped)
Uncertainty at runtime1.96 × the standard deviation of per-tree predictions, at least 1 cycle: a transparent proxy for a prediction interval
Artifact size161 MB (the largest of the four checkpoints)
The RUL model in one table. Source: scripts/train_rul.py, src/machina_harness/rul.py
Remaining-useful-life error (cycles)Remaining-useful-life error (cycles)0.0020.0040.0060.0080.000.791.39Engine-level holdout(20 engines from the training fleet)62.3671.72Official FD001 test fleet(100 engines)MAERMSE
Figure 10 Two numbers, one honest. A 0.79-cycle error on the training fleet’s holdout is far too good to be real; the 62.4-cycle MAE on the official test fleet is the figure that counts, and the repository calls it “a starting point for model improvement, not a production-quality result.” A likely reason for the gap, inferred from the training code and not stated in the repository: the cycle_fraction feature is cycle divided by the engine’s total life during training, which leaks the answer, while at test time it is computed from the truncated observed trajectory. Source: github.com/Clerktree/machina-intelligence · artifacts/rul-cmapss/metadata.json; docs/RUL_BASELINE.md

Process quality: accuracy 98%, macro-F1 55%

The third capability predicts a process failure mode from six machine settings, using the synthetic UCI AI4I 2020 dataset: tool wear, heat dissipation, power, overstrain or random failure, as a probability that something other than “normal” is happening.

Per-class F1, process-failure model (AI4I 2020, synthetic)Per-class F1, process-failure model (AI4I 2020, synthetic)normal2,413 windows · 96.5% of the data0.989overstrain190.800heat dissipation290.750power230.651tool wear120.111random40.000
Figure 11 Accuracy 0.978 looks excellent; macro-F1 0.550 tells the real story. 96.5% of the data is “normal”, so the headline number is mostly one easy class, while tool-wear recall is 8.3% and the “random” failure mode is never found. The repository calls the model “a contract and integration baseline rather than a deployment claim.” Source: github.com/Clerktree/machina-intelligence · artifacts/ai4i-quality/metadata.json
Row-normalised confusion (percent of each true class)Row-normalised confusion (percent of each true class)heatnormaloverpowerrandomtoolheat dissip.72280000normal0990000overstrain0595000power03096100random01000000tool wear0838008
Figure 12 Where the failures go. Each row is a true failure mode; the cell is the share predicted as each column. Most missed failures are called “normal”: 10 of 12 tool-wear windows and all four “random” failures. Row-normalising makes the minority classes legible; in raw counts the normal cell (2,391) would swamp everything. Source: github.com/Clerktree/machina-intelligence · artifacts/ai4i-quality/metadata.json; derived percentages

The agent: routing, not reasoning from scratch

Machina’s design splits intelligence in two. Specialist models operate on machine signals; a small language model translates a natural-language request into the right tool call. The agent is a Mistral-7B-Instruct adapter, trained with QLoRA to emit one structured tool call.

The tool-routing agentA user request goes to a Mistral-7B agent with a LoRA adapter, which emits a tool call; the MCP server runs the specialist models; the tool result goes back to the agent for a grounded answer.Usernatural languageMachina agentMistral-7B-Instruct-v0.3+ LoRA adapter[TOOL_CALLS]JSON: tool + argumentsMCP server15 tools · specialist modelstool result → final grounded answerThe request-execute-answer loop is described in the docs; the executor itself is left to the deployer.Specialist models read signals; the agent only orchestrates. It is a fine-tune, not a foundation model.
Figure 13 The reasoning interface. The small language model is the orchestrator, not a replacement for the signal models: it picks the smallest useful tool, and its system prompt tells it to inspect before acting, never invent sensor readings and separate observed tool results from hypotheses. Source: github.com/Clerktree/machina-intelligence · src/machina_harness/agent.py; docs/MACHINA_AGENT_TRAINING.md
SettingValue
Base modelmistralai/Mistral-7B-Instruct-v0.3 (Apache-2.0, native function calling)
MethodQLoRA: 4-bit NF4, double quantisation, bf16 compute
LoRArank 32 · alpha 64 · dropout 0.05 · q, k, v, o, gate, up, down projections
Optimisationlearning rate 2e-4, cosine, warm-up 5% · batch 1 × accumulation 8 · 2 epochs · paged AdamW 8-bit
Sequence length1,280 tokens, gradient checkpointing
Lossonly on the first assistant tool-call turn: the prompt and tool schema are masked
Data1,200 synthetic examples: 10 tools × 120; 92 / 8 train / eval split
Hardwareone lab GPU (RTX 4500 Ada)
The fine-tuning recipe. The adapter is small and trained to do one thing: route. Source: scripts/train_machina_agent.py, docs/MACHINA_AGENT_TRAINING.md

The training set is a synthetic, grounded routing set: it teaches routing and response style, not machine facts. It contains 1,200 examples, exactly 120 for each of ten tools, built from 50 distinct user strings. That makes it a good fit for a router and a poor basis for claims about generalisation, which is why the repository says to validate the adapter on held-out routing examples before calling it production-ready.

The 15 MCP tools, and the 10 the router is trained on
MCP toolIn the router’s training schema
machine_harness_healthyes
list_machine_intelligence_capabilitiesyes
machina_platform_snapshotyes
list_machine_assetsyes
list_registered_modelsyes
search_machine_knowledgeyes
prepare_maintenance_briefyes
estimate_remaining_useful_lifeyes
analyze_machine_energyyes
predict_process_qualityyes
analyze_machine_windowno
register_machine_assetno
record_machine_telemetryno
record_maintenance_eventno
index_maintenance_documentno
The 15 MCP tools and which of them the agent was trained to route to. Fault diagnosis (analyze_machine_window) and the write tools are not in the trained schema yet. Source: src/machina_harness/mcp_server.py, scripts/build_agent_dataset.py

Evidence for the language layer

The “reasoning” side of the platform is a retrieval brief. Maintenance documents are indexed in an SQLite full-text table and searched with BM25; prepare_maintenance_brief returns the asset record, the top evidence passages, the registered model ids and three fixed instructions for whatever language model consumes them:

copilot.py · fixed grounding instructions returned with every maintenance brief
Use only the returned evidence and model outputs as factual support.
Separate observed signals from hypotheses.
Recommend qualified human inspection for safety-critical decisions.

Governance as code

The least glamorous part of the repository is the part most worth copying. The platform treats every analysis as something that may be audited, challenged and overridden.

Lifecycle of one analysis requestA request passes API-key check, validation, inference and the confidence gate, returns a finding that always requires human review, and writes a metadata-only audit row.RequestPOST/v1/analyzeAPI keyX-Machina-API-KeyValidatefinite · ≥ 3samplesInferrules +classifierGateconfidence≥ 0.65Respondhuman reviewrequiredinference_audits rowaudit_id · request_id · machine_id · timestamp · status · model_versionanomaly_score · predicted_fault · fault_confidence · abstained · sensor sample counts · latency_mslogged
Figure 14 A governed request. The audit row is deliberately metadata-only: it keeps sample counts and scores, never the raw sensor values, so the log is useful for traceability without becoming a data-leak surface. X-Request-ID is echoed or generated on every response and logged with latency. Source: github.com/Clerktree/machina-intelligence · src/machina_harness/api.py, platform.py
ControlHow
Authenticationoptional MACHINA_API_KEY; when set, every route except /health needs X-Machina-API-Key (401 otherwise)
Deploymentthe compose file refuses to start without the key, runs read-only with a tmpfs /tmp, no-new-privileges, all capabilities dropped, health-checked every 30 s
Abstentionfault confidence below MACHINA_FAULT_MIN_CONFIDENCE (default 0.65) returns no diagnosis
Human in the loophuman_review_required = True on every finding; critical findings advise stopping reliance on the automated diagnosis
TraceabilitySHA-256 of every model artifact on /v1/model-health; /ready reports “degraded” if a plugin is missing
Release gateverify_release.py requires README, metadata and weights per artifact and, for the enhanced bearing model, a leave-one-speed-out floor and a “not evidence of generalization” sentence in its model card
Governance controls, all implemented in code rather than promised in prose. Source: src/machina_harness/api.py, docker-compose.yml, scripts/verify_release.py

Portability and the edge

“Portable” here has a concrete meaning: the models are plain scikit-learn joblib files with a metadata JSON and a model card, they are registered as plugins that can be swapped, and path overrides let a deployment point to its own checkpoints. They run on a CPU with no network access at inference.

Size of each model artifact bundled in the Docker image (MB)Size of each model artifact bundled in the Docker image (MB)RUL · ExtraTrees (C-MAPSS)161.2 MBProcess quality · ExtraTrees31.2 MBBearing · calibrated ET (shipped)6.4 MBBearing · baseline RF6.4 MB
Figure 15 All four artifacts total about 205 MB, dominated by the RUL forest, and run on a CPU with scikit-learn and no network calls at inference. That is the extent of the edge story today: the docs describe merging the agent adapter and exporting a quantised GGUF or AWQ artifact for local runtimes, but that has not been done, and no latency, memory or hardware benchmarks are published. The base image is python:3.12-slim. Source: github.com/Clerktree/machina-intelligence · artifacts/*/model.joblib (Git LFS pointer sizes), Dockerfile

Run it

Python 3.10 or newer. Fetch the large files with Git LFS, install the package and start the API; the interactive documentation is at /docs.

shell
git clone https://github.com/Clerktree/machina-intelligence && cd machina-intelligence
git lfs pull                                   # model.joblib files are Git LFS objects
python3 -m venv .venv && source .venv/bin/activate
pip install -e '.[dev]'                        # extras: dev, train, mcp, runtime, agent
uvicorn machina_harness.api:app --reload       # http://127.0.0.1:8000/docs

curl -X POST http://127.0.0.1:8000/v1/analyze \
  -H 'content-type: application/json' -d @configs/sample-window.json
curl http://127.0.0.1:8000/v1/model-health

# container (the key is mandatory)
export MACHINA_API_KEY="replace-with-a-long-random-secret"
docker compose up --build -d
curl http://127.0.0.1:8000/ready -H "X-Machina-API-Key: $MACHINA_API_KEY"

# MCP over stdio (e.g. for Grok Build)
pip install -e '.[mcp]' && python3 -m machina_harness.mcp_server

A few practical notes. The classifier imports scipy.signal.hilbert, which arrives through scikit-learn rather than being declared on its own. The model cards declare Apache-2.0. The capability registry is the source of truth for what is live: fault diagnosis, anomaly detection, remaining useful life, energy intelligence and quality prediction are available, while the maintenance copilot and machine knowledge capabilities are marked as planned.

The decision-support framing is repeated in the repository, and it is worth repeating here: the model is decision support, not a safety controller. Human inspection and site-specific validation are required before any maintenance action.

All articles

Research and building cool stuff

© 2026, Animesh Mishra

GitHub|LinkedIn

New Delhi, India