Turnkey AI coding assistant that runs entirely on-premises — tab completions, chat, and inline diffs powered by open-source models with zero internet dependency.
Cloud-based AI coding tools are off-limits in secure environments. Your developers watch the productivity revolution from the sidelines.
Developers in air-gapped environments miss out on 30–40% productivity gains that cloud-connected teams enjoy from AI coding tools.
78% of developers in secure environments say they want AI coding assistance but are blocked by data exfiltration policies.
Self-hosting open-source AI tools means assembling models, inference servers, authentication, and monitoring yourself — then maintaining all of it.
A self-contained AI coding platform that deploys behind your firewall. Open-source models, zero external dependencies, full organizational control.
Live walkthrough, GPU sizing, security review — 30 minutes, no sales pitch.
Nine core capabilities that bring the full AI coding experience to your air-gapped network.
Context-aware code suggestions as you type, powered by a purpose-built fill-in-the-middle model.
Ask questions about your codebase, generate tests, refactor code — all inside your IDE.
Review AI-suggested changes inline with accept/reject controls. No context switching.
Monitor usage, manage seats, configure models, and view analytics across your organization.
Single sign-on via your existing directory service. No separate accounts to manage.
IDE plugins reach the M8 server over HTTPS (TLS 1.2+), on by default. Upload your own certificate or use the one generated at install.
Per-user rate limits prevent abuse and keep GPU capacity fairly shared across your team.
Prometheus-compatible metrics endpoint. Grafana dashboards included out of the box.
Docker bundle with pre-baked model weights. No internet required at any point during deployment.
Download the M8 Docker bundle with pre-baked model weights on a connected machine. Transfer to a portable drive.
Carry the bundle across the air gap. Load Docker images on your GPU server inside the secure network.
Run the guided installer. It detects your GPU, configures vLLM and the API server, connects to your directory, and verifies the install end to end.
Both models are Apache 2.0 licensed with no usage restrictions. Developed by organizations in NATO-allied countries.
Purpose-built for inline code completion. Trained on permissively-licensed code. Optimized for sub-200ms low-latency suggestions.
State-of-the-art open-weight model for code understanding, generation, refactoring, and natural language interaction with large context windows.
M8 scales from a single RTX 4090 supporting a small team to NVIDIA B300 clusters serving hundreds of developers.
| GPU | VRAM | User Capacity | Price Point |
|---|---|---|---|
| RTX 4090 | 24 GB | 5–15 users | ~$2K GPU |
| L4 | 24 GB | 5–15 users | ~$0.85/hr |
| 2–4x L4 | 48–96 GB | 30–80 users | ~$3–5/hr |
| L40S | 48 GB | 30–60 users | ~$1/hr |
| A100 40GB | 40 GB | 30–50 users | ~$2/hr |
| A100 80GBRecommended | 80 GB | 50–100 users | ~$3/hr |
| H100 | 80 GB | 100–200 users | ~$4/hr |
| B300 | 192 GB | 300–500 users | Enterprise |
See how M8 compares to cloud-based alternatives on the dimensions that matter most to secure environments.
M8 is the only turnkey AI coding assistant that deploys behind an air gap with open-source models.
| Feature | M8 | GitHub Copilot | Cursor | Tabnine Enterprise | Continue.dev |
|---|---|---|---|---|---|
| Air-gapped deployment | Turnkey | Cloud only | Cloud only | Proprietary | DIY |
| Open-source models | Apache 2.0 | Proprietary | Proprietary | Proprietary | Yes |
| On-premises | Full stack | No | No | Limited | Model only |
| Tab completion | Yes | Yes | Yes | Yes | Yes |
| Chat assistant | Yes | Yes | Yes | No | Yes |
| Inline diffs | Yes | Yes | Yes | No | Limited |
| Admin dashboard | Full | No | No | Yes | No |
| LDAP/AD SSO | Yes | No | No | Yes | No |
| GPU profiles | 8 pre-configured | N/A | N/A | N/A | No |
| Monitoring (Grafana) | Pre-built | N/A | N/A | Yes | No |
M8 runs entirely on hardware you own. Models, inference, and data stay inside your network, and nothing depends on an outside service. Every control below is built into the product today, so your security team can verify it for themselves.
Everything you need to know about M8.
All models run locally via vLLM on your GPU server. The air-gap bundle ships the Docker images together with pre-baked model weights, so neither installation nor day-to-day operation needs an internet connection.
M8 uses purpose-built open-source models — one optimized for fast tab completion (fill-in-the-middle), another for chat and code generation. Both are Apache 2.0 licensed, sourced from allied-nation research institutions.
One guided installer run on your GPU server. It detects the GPU, loads the bundled images and model weights, configures vLLM, the API server, and directory sign-in, then verifies the install and diagnoses anything that fails.
M8 supports 8 GPU profiles ranging from RTX 4090 (5–15 users) to NVIDIA B300 (300–500 users). See the GPU compatibility table above for full details.
M8 is built for networks that cannot send code to the cloud: everything runs on your own hardware with no internet dependency, with directory sign-in, HTTPS, hashed API keys, and an exportable audit log built in. Whether a deployment meets a specific compliance framework depends on your environment and controls, so we walk your security team through the architecture in detail during the demo.
GitHub Copilot requires internet connectivity and sends code to Microsoft servers for processing. M8 runs entirely on-premises with open-source models — similar tab completion and chat features, but zero data ever leaves your network.
M8 runs on open-weight models licensed under Apache 2.0, served by vLLM on your own GPUs. Usage data, audit logs, and configuration are stored in PostgreSQL on your own servers — not with us.
Yes. Request a demo and our engineers will walk you through a live air-gapped deployment, review your security architecture, and size GPU hardware for your team.