Content
# MagLab Studio
<p align="center">
<strong>A local-first scientific workflow control plane for choosing the model, harness, evidence policy, and workflow independently.</strong>
</p>
<p align="center">
<a href="README.md">English</a> ·
<a href="README.ko.md">한국어</a> ·
<a href="README.ja.md">日本語</a> ·
<a href="README.zh-CN.md">简体中文</a>
</p>
<p align="center">
<a href="https://github.com/TaewoooPark/MagLab-Studio/releases/latest"><img alt="Latest release" src="https://img.shields.io/github/v/release/TaewoooPark/MagLab-Studio?style=flat-square&color=111111"></a>
<a href="https://github.com/TaewoooPark/MagLab-Studio/actions/workflows/quality.yml"><img alt="Quality gates" src="https://img.shields.io/github/actions/workflow/status/TaewoooPark/MagLab-Studio/quality.yml?branch=main&style=flat-square&label=quality&color=111111"></a>
<a href="LICENSE"><img alt="Apache-2.0 license" src="https://img.shields.io/badge/license-Apache--2.0-111111?style=flat-square"></a>
<img alt="Desktop platforms" src="https://img.shields.io/badge/desktop-macOS%20%7C%20Linux%20%7C%20Windows-111111?style=flat-square">
<img alt="Local first" src="https://img.shields.io/badge/runtime-local--first-111111?style=flat-square">
</p>

<p align="center"><sub>Concept illustration — the screenshots below are captured from the running application.</sub></p>
MagLab Studio separates four decisions that are usually fused together:
1. **Model** — the intelligence and provider protocol.
2. **Harness** — the agent loop, tools, permissions, and session semantics.
3. **Capability bundle** — skills, plugins, agents, MCP servers, and hooks.
4. **Research policy** — workspace scope, approvals, validation, provenance, and replay.
The result is not another chat shell. It is a desktop control plane where a route such as **GPT through Claude Code**, **Claude through OpenCode**, or a local model through a compatible harness can be selected explicitly, tested against its real transport, and attached to scientific operations that do not depend on the model.
> [!IMPORTANT]
> Version `0.2.0` adds the first four gated Research OS milestones: authenticated local review, immutable Research Objects with stale propagation, and a reviewable FMR narrative-to-FitSpec proposal workflow. This is proposal and provenance infrastructure, not a validated FMR fit. The Week 4 evidence is published in the [Research OS manifest](quality/research-os/evidence/week-04.json); later data, fit, validator, claim, replay, and baseline gates remain fail-visible.
>
> Local review records identify an authenticated desktop session, not a person's civil identity or expertise. Each decision requires a server-issued, one-time authorization. External human and physicist validation remain `not-run` unless a release evidence manifest says otherwise.
## See it running
### Pick a harness and a model independently

The launchpad reads the versioned harness registry, detects installed runtimes, loads the provider-neutral model catalog, and shows native, managed-install, and bridged routes without making the user scroll through a chat first.
### Run deterministic, cited science outside the model

The screenshot above is a real `physics.compute` execution for magnetic exchange length. The result carries the operation version, cited formula, validated finite inputs, Python runtime identity, and stable input/result hashes.
### Review and execute a typed FMR fit without turning chat into state
The Research surface accepts an exact FMR narrative plus bounded literal dataset
headers and sample strings—not the original or raw dataset bytes. It compiles
alternative parses and unresolved spans for an exact one-time-authorized review.
Acceptance materializes immutable Question, Hypothesis, and narrative FitSpec
objects that remain in proposal state; it does not ingest data or execute a fit.
Execution starts only after a separately prepared strict
`maglab-fmr-fit-spec/v1` has been trusted-accepted and bound to an immutable
DatasetVersion. **FMR fit review** loads those exact ResearchIR refs, displays
equations, applicability, units, parameter bounds, weighting, solver, and five
required sensitivity analyses, and executes the pinned deterministic kernel
only after another one-time-authorized desktop review. Before showing scientific
result details, the desktop reads back the exact FitResult and audited
DatasetView projection through separate ResearchIR/data-lineage routes. This
read-back verifies refs and lineage; it is not an independent scientific
validator. A failed read-back suppresses result details, locks re-execution, and
retains the durable FitResult ref for recovery.
This development path is limited to the frozen tabular field-swept FMR vertical.
The NIST publisher-processed observables remain scientifically `inconclusive`:
raw complex-S21 extraction, background/phase choice, mode assignment,
calibration covariance, and independent replicate variation were not
reproduced. Independent numerical validation, blind/holdout scoring, interval
coverage, claim promotion, external human or physicist validation, Week 7 fit
clean replay 3/3, simulation, imaging, magnetic texture, and orbitronics remain
`not-run`; non-FMR verticals are `unsupported`.
### Prepare editable research artifacts with an independent gate

Poster, deck, journal-figure, and manuscript jobs freeze their source before a managed authoring engine receives it. The producing engine cannot certify its own output: separate validators inspect source identity, native editability, geometry, deliverables, and integrity.
### Compare routes using evidence, not branding

Evidence Lab distinguishes registry declarations, executable detection, transport connectivity, behavioral certification, and run evidence. Recommendations expose their factors instead of hiding them behind a single opaque score.
### Keep physical actions behind a local review boundary

Instrument plans use typed SiLA identifiers, finite parameter bounds, a simulated safety check, live endpoint attestation, one-time local desktop review authorization, ephemeral mTLS material, and a separate emergency-stop lane. The simulated check covers declared schema, bounds, order, and state only. This is an orchestration boundary, not a laboratory safety certification or external human validation.
## What is included
| Plane | What it provides | Evidence boundary |
| --- | --- | --- |
| Universal router | Model/harness-independent sessions over app-server, ACP, structured CLI, text CLI, and local protocol bridges | Declared, detected, connected, and certified states remain distinct |
| Capability plane | Shared skills, plugins, agents, MCP servers, instructions, and lifecycle hooks | Executable capabilities are opt-in; unsupported delivery is marked degraded |
| Execution ledger | Append-only runs, redacted events, ATIF trajectories, OpenTelemetry GenAI traces | Records observed behavior; does not infer scientific truth |
| Science Kernel | Versioned formulas, unit conversion, physical-range checks, strict schemas, stable hashes | Deterministic operations are separate from model prose |
| Research OS | Authenticated review, immutable versioned Research Objects and DatasetVersions, event replay, stale propagation, proposal-state narrative formalization, and separately accepted typed field-swept FMR execution | NIST publisher-processed observables remain `inconclusive`; independent validation, blind metrics, claims, Week 7 fit clean replay 3/3, and non-FMR verticals remain `not-run` or `unsupported` |
| Research evidence | Frozen source bytes, exact quote locators, local session review, integrity events | Model judgments stay proposed until a one-time-authorized desktop review; reviewer identity and expertise are not verified |
| Artifact plane | Posters, decks, journal figures, manuscripts, independent OOXML/PDF/ZIP validation | Generated output is not complete until the independent gate passes |
| Research Capsules | RO-Crate 1.3, Process Run Crate 0.5, BagIt 1.0, byte-level ZIP reopen checks | Integrity and replay metadata do not imply authenticity or correctness |
| Experiment/ELN | Immutable raw-data hashes, assay review, ISA-JSON, ELN exchange | Preserves records and lineage; it does not validate experimental design |
| HPC | Frozen Slurm REST plans, reviewed inputs, ephemeral credentials | Submission requires a one-time-authorized local review and a live cluster profile; real-world operator identity is outside this release |
| Instruments | Typed bounded plans, simulated safety check, live SiLA attestation, local review, emergency-stop lane | Simulation is not device validation; real devices still require institutional and vendor safety procedures |
## Supported harnesses
The versioned registry currently includes:
| | | | |
| --- | --- | --- | --- |
| Codex | Claude Code | OpenCode | Pi |
| Qwen Code | OpenHands | Goose | Kimi Code |
| Cline | Crush | Aider | Gemini CLI |
| OpenClaw | | | |
Existing installations are detected automatically. Missing runtimes can be installed into `~/.maglab-studio/runtimes` without modifying the global npm or Python environment. Authentication remains in each harness's own account store.
The managed Apache-2.0 [Bifrost](https://github.com/maximhq/bifrost) gateway can connect OpenAI, Anthropic, Gemini, OpenRouter, Mistral, Groq, xAI, Ollama, vLLM, SGLang, and custom OpenAI-compatible endpoints. OpenRouter is optional; local and direct provider endpoints are first-class routes.
Model/harness independence is a protocol claim, not a parity promise. A non-native model remains selectable when the protocol can be translated, but the UI marks missing tool calling, reasoning, multimodality, or session behavior instead of inventing support.
## Start here
### Install the release
Download the newest package from [GitHub Releases](https://github.com/TaewoooPark/MagLab-Studio/releases/latest).
The `0.2.0` macOS package is unsigned and unnotarized. Verify the release asset and use the source build if your environment requires a trusted signing chain. Linux and Windows source builds are checked in CI; packaged installers will be added release by release.
### Run from source
Requirements:
- Node.js 22 or newer
- Rust stable
- The [Tauri 2 prerequisites](https://v2.tauri.app/start/prerequisites/) for your operating system
- `uv` for the complete release gate; the app can provision a pinned managed `uv` runtime for Science Kernel use
- A login or API credential only for the providers and harnesses you choose
```bash
git clone https://github.com/TaewoooPark/MagLab-Studio.git
cd MagLab-Studio
./run-maglab-studio.sh
```
The launcher installs the locked desktop JavaScript dependencies on first run, then starts the Tauri application. Managed runtimes and authoring engines are stored outside the repository under `~/.maglab-studio`.
### Build a production desktop bundle
```bash
cd apps/desktop
npm ci
npm run desktop:build
```
### Run the complete release gate
```bash
node scripts/release-gate.mjs
```
The gate executes 13 blocking checks across the router, Claude host, React production build, three locked Python services, and the Rust host. It is network-independent and hardware-independent by design, so mocked Slurm or simulated instruments are never promoted into live-conformance claims.
## First session in five steps
1. Open **New chat**.
2. Pick the harness on the left and model maker/version on the right.
3. Open **Shared capability bundle** and enable only the skills, agents, plugins, MCP servers, and hooks you trust for this session.
4. Confirm the workspace and permission profile.
5. Continue to the prompt, review every external or high-risk action, and use **Evidence Lab** or **Claims & evidence** when the result must be auditable.
For setup, routing, evidence review, science operations, artifact creation, HPC, instruments, troubleshooting, and safety notes, read the full [English user manual](docs/manuals/en/index.md).
Manuals: [English](docs/manuals/en/index.md) · [한국어](docs/manuals/ko/index.md) · [日本語](docs/manuals/ja/index.md) · [简体中文](docs/manuals/zh-CN/index.md)
## Architecture
```mermaid
flowchart LR
U["Desktop workspace"] --> P["Session policy"]
P --> H["Harness adapter"]
P --> C["Capability bundle"]
M["Provider model"] --> G["Loopback Bifrost gateway"]
G --> H
C --> H
H --> L["Append-only execution ledger"]
H --> S["Deterministic Science Kernel"]
H --> A["Artifact runtime"]
H --> R["Research evidence ledger"]
H --> X["HPC / instrument boundary"]
S --> V["Independent validation"]
A --> V
R --> V
L --> E["Evidence Lab"]
V --> E
E --> K["Verifiable Research Capsule"]
```
### The separation that matters
- **Harness adapters** own session semantics, prompt delivery, native tools, and runtime-specific events.
- **The gateway** translates provider protocols on loopback and keeps provider credentials out of arguments and client-side state.
- **The capability compiler** snapshots selected cross-harness resources into a content-addressed session bundle.
- **The evidence plane** records observed execution separately from declarations.
- **The science and artifact planes** are callable from the desktop and MCP-capable harnesses through the same typed contracts.
- **Research Capsules** package only explicit inputs, accepted evidence, validated artifacts, and replay metadata.
## Repository map
```text
apps/desktop/ React 19 interface and native Tauri 2 host
services/router/ Universal routing, capabilities, evidence, HPC, instruments
services/claude-host/ Claude Code ACP integration boundary
services/science-kernel/ Typed deterministic scientific operations
services/artifact-kernel/ Journal figures and evidence-bound document rendering
services/sila-sidecar/ Lockfile-pinned, fail-closed SiLA 2 device boundary
harnesses/registry.json Versioned harness and capability registry
quality/ Parity inventory, pins, gates, and release conformance
docs/ Architecture, research, release, and multilingual manuals
```
## Data, credentials, and safety
- Provider secrets are stored in the operating-system keychain or native harness account store.
- Gateway traffic is loopback-only by default.
- Workspace evidence stays in the selected workspace; managed runtimes live under `~/.maglab-studio`.
- Research Capsules include only explicitly selected files and accepted evidence.
- Secrets are redacted from the append-only execution ledger.
- Hooks, executable plugins, MCP servers, unrestricted execution, HPC submission, and instrument execution require explicit authority.
Harnesses still execute code with the permissions granted to their local processes. Prefer workspace-write mode, keep approvals enabled, and add an external sandbox for runtimes that do not enforce one themselves.
Real HPC clusters and laboratory instruments require local risk assessment, trained operators, calibration and maintenance records, access controls, hardware-in-the-loop testing, and vendor emergency procedures. Cancelling a network request is never proof that physical equipment stopped.
## MagLab relationship and source independence
[MagLab](https://github.com/TaewoooPark/MagLab) is the CLI reference whose public capability inventory defines an important parity target. MagLab Studio is not a skin over that repository and does not copy its implementation. It is an independently authored desktop and routing system with its own protocol boundaries, validation gates, scientific runtime, evidence model, and product documentation.
The frozen inventory currently covers 116 MagLab CLI leaves and 42 provider-neutral tools. Inventory is not completion: each surface advances only when its result is editable where relevant, independently validated, provenance-complete, and replayable. See the [integration strategy](docs/maglab-integration-strategy.ko.md), [beyond-parity research](docs/beyond-parity-research.ko.md), and [current release boundary](docs/releases/v0.2.0.md).
## Development and verification
Useful focused commands:
```bash
npm ci --prefix apps/desktop
npm ci --prefix services/router
npm ci --prefix services/claude-host
npm test --prefix services/router
npm test --prefix services/claude-host
npm run build --prefix apps/desktop
uv run --project services/science-kernel --frozen pytest
uv run --project services/artifact-kernel --frozen pytest
uv run --project services/sila-sidecar --frozen pytest
cargo test --locked --manifest-path apps/desktop/src-tauri/Cargo.toml
```
CI repeats the complete gate on Ubuntu and checks the native host and production interface on macOS and Windows. Lockfiles are enforced, and third-party GitHub Actions are pinned to immutable revisions.
Before opening a pull request:
1. Add or update the contract test for the changed boundary.
2. Keep unsupported behavior fail-visible.
3. Update the conformance record or manual when the assurance boundary changes.
4. Run `node scripts/release-gate.mjs`.
## Documentation
- [English user manual](docs/manuals/en/index.md)
- [Evidence Lab](docs/evidence-lab.md)
- [Science Kernel](docs/science-kernel.md)
- [Research evidence](docs/research-evidence.md)
- [Research Capsules](docs/research-capsules.md)
- [Artifact runtime](docs/artifact-runtime.md)
- [Experiment ledger](docs/experiment-ledger.md)
- [HPC jobs](docs/hpc-jobs.md)
- [Instrument safety](docs/instrument-safety.md)
- [Capability compatibility](docs/capability-compatibility.md)
- [Capability routing research](docs/capability-routing-research.md)
- [Release 0.2.0 assurance boundary](docs/releases/v0.2.0.md)
## Versioning
The project uses semantic versioning while pre-1.0:
- **Patch** — compatible fixes, adapters, and presentation improvements.
- **Minor** — new router capabilities or material runtime expansion.
- **Major** — incompatible protocol or security-model changes.
## License
MagLab Studio is licensed under [Apache License 2.0](LICENSE). External harnesses remain separate programs under their own licenses. See [THIRD_PARTY_NOTICES.md](THIRD_PARTY_NOTICES.md) for attribution and compatibility notices.
Connection Info
You Might Also Like
everything-claude-code
Complete Claude Code configuration collection - agents, skills, hooks,...
markitdown
MarkItDown-MCP is a lightweight server for converting URIs to Markdown.
cc-switch
All-in-One Assistant for Claude Code, Codex & Gemini CLI across platforms.
servers
Model Context Protocol Servers
servers
Model Context Protocol Servers
Time
A Model Context Protocol server for time and timezone conversions.