DWG NO. HC-0.8.0 CLI REV. 0.7.0 SCALE LOCAL = CLOUD LICENSE PolyForm Noncommercial 1.0 TELEMETRY ANONYMOUS, OPT-OUT

Your code doesn’t
leave the building.

Heap Code is a model-agnostic AI coding assistant for VS Code, your terminal, and your browser — chat, completions, inline edit, an autonomous agent, and AI PR review. Point it at Ollama on your own machine, or bring your own key to OpenAI, Groq, or any OpenAI-compatible endpoint. No account. No proprietary backend.

Extension v0.8.0 · CLI v0.7.0 · free for personal, internal & noncommercial use

Diagram of where your code travels depending on local or cloud mode YOUR MACHINE VS Code Heap Code Ollama / LM Studio OpenAI / Groq / …

Local: prompts, completions, and embeddings run entirely inside the boundary above — nothing crosses the line.

The terminal CLI

heapcode — zsh
npm install -g @heaplabs/heapcode-cli
heapcode
 
> add rate limiting to the upload route
 
⏺ read_file src/routes/upload.ts
⏺ edit_file src/routes/upload.ts
⏺ run_tests pnpm test
✓ 48 passed
 
heapcode -p "fix the failing test" --json

Node.js 20+ · macOS, Linux, Windows · no account, no sign-in

  • Same engine as the extension. The agent, permissions, checkpoints, semantic search, and MCP support are shared code — not a reduced second implementation.
  • /pr-review reviews your branch's PR file by file, then posts a real GitHub review with inline comments and one-click suggestions — only after you approve it. /pr-review deep adds a verification pass that drops false positives first.
  • Headless mode. -p "task" runs the whole agent loop with no TTY — plain text or --json events, four permission modes for unattended runs, non-zero exit on failure. Built for CI.
  • Every action gated. Same permission prompts as the editor, plus shadow-git checkpoints you can rewind step by step.
  • Your models. Local Ollama or LM Studio, a box on your LAN, or any OpenAI-compatible endpoint — configured once, per role.

The browser UI

127.0.0.1:7411
youwhy is the upload route timing out?
⏺ semantic_search upload timeout
⏺ read_file src/routes/upload.ts
⏺ edit_file src/routes/upload.ts  +12 −3
The multipart parser had no size cap, so a large file…

Runs on your machine · localhost by default · no account, no hosted service

  • One command. heapcode web in any folder opens a browser tab pointed at that workspace — same agent, same permissions, same checkpoints as the CLI and the extension.
  • A real workspace panel. Per-file diffs, one-click revert, a checkpoint timeline to rewind to, the file tree, streamed command output, and what the agent has actually indexed about your code.
  • Things a terminal can’t do. Paste a screenshot straight into the message box. Get a desktop notification when a long run finishes in a background tab. Read rendered diagrams and HTML the agent built for you.
  • Close the tab, keep the run. The server owns the work, not the browser — reopening reattaches to whatever is still going.
  • Localhost by design. A fresh token per launch, never left in the address bar. --host opts in to reaching it from your phone, and says so loudly. Your API keys stay on the server and never reach the page.

New — Heap Chat

127.0.0.1:7411/chat
youwhat did I pay for the boiler service last year?
⏺ semantic_search boiler service invoice
⏺ read_file Home/2025/invoice-0412.pdf
£184.00, paid 12 April 2025 — from invoice-0412.pdf, line 9.

Ships in the CLI · heapcode chat [folder] · or at /chat under heapcode web

  • Your files, not your code. Point it at a folder of PDFs, Word documents, spreadsheets and photos. It reads them, searches them, and answers from them — a scanned receipt is findable by what it says.
  • Every answer shows its source. It says which file an answer came from, and a figure can be traced back to the line it was read from.
  • It never writes over your files. A summary, a table or a report lands beside the conversation as a draft. Saving it into your folder is your call.
  • Same engine, its own product. Its own tools, prompt, history and memory — sharing only the engine, the models you configured, and the index with Heap Code.
  • Connectors with sign-in. MCP servers work here too, including hosted ones that need you to log in — and it asks before running one.

Also on this engine — heapbrowse

  • A Chrome side panel that works the page you are on. Ask about it, or have it search, filter, compare across pages and fill in forms — and hand back the steps only you should take.
  • Nothing happens without you seeing it. Every click or form fill is checked against what the element really is, and asks first. Banking and password sites are refused outright.
  • Your own models. Point it at Ollama on your machine or your LAN — no server setup needed — and the pages you read never leave home.

Chrome extension · side panel · sites granted one at a time. Not on the Chrome Web Store yet: download the heapbrowse .zip from the latest release, unzip it, and choose “Load unpacked” in chrome://extensions.

Spec sheet

Telemetry Anonymous usage events only (feature usage, error counts) — never code, prompts, or file contents/paths. Opt-out via a setting.
Account required No. No sign-in, no license server.
Model lock-in None — any OpenAI-compatible endpoint, switch providers per-role (chat / edit / apply / completion / agent / embeddings / rerank).
Local model support Yes — Ollama, LM Studio, vLLM, LocalAI.
Surfaces VS Code extension, a standalone terminal CLI (heapcode), and a local browser UI (heapcode web) — one shared engine, so behaviour matches. Chat history, connections and project MCP servers are shared between them. Same engine: Heap Chat (heapcode chat) and the heapbrowse Chrome extension.
API key storage Extension: OS keychain via VS Code SecretStorage. CLI: a chmod 600 file in ~/.heapcode — no keychain dependency, so headless and CI machines work. Never in settings files either way.
License / price Free for personal, internal, research, nonprofit, education, and government use (PolyForm Noncommercial 1.0).

Parts list — works with

Localruns on your machine

  • Ollama
  • LM Studio
  • vLLM
  • LocalAI

Cloudbring your own key

  • OpenAI
  • Azure OpenAI
  • OpenRouter
  • Groq
  • Together AI
  • NVIDIA NIM

Features

Chat

Streaming markdown, slash commands, @file/@selection/@workspace context, per-workspace history.

Completions

Ghost-text FIM tuned per model family, debounced and cancellable, repo-aware via the semantic index.

Inline edit

Select code, describe the change, review a native diff, accept from the title bar.

Agent mode

Reads, searches, edits, and runs commands autonomously. Shift+Tab cycles how much runs unattended — Plan (read-only, mutating tools aren’t even offered), Confirm, Auto-edit, or Auto — changeable mid-run, and Auto still stops for anything destructive. One-click revert on everything.

Semantic search

AST-aware chunking, hybrid embeddings + keyword search, degrades gracefully with no embedder configured.

MCP & tool interop

Register Model Context Protocol servers, including hosted ones that need a sign-in; each one gets only the environment variables it needs, not your whole shell. Other extensions' language-model tools show up in agent mode too.

Safety guardrails

Untrusted content — files, URLs, search results, MCP output — is tagged and quarantined before it reaches the model; hallucinated package installs are checked against the registry and blocked automatically; and agent web fetches can't reach private, loopback, or cloud-metadata addresses, at any redirect hop.

Granular checkpoints

Every tool call gets its own shadow-git snapshot — rewind any single step, not just the whole turn or the whole session.

Personas & Plan/Act

Architect, Debug, and Reviewer personas restrict which tools the agent is even offered; an optional plan-then-approve gate holds every change until you say go. Personas and permission modes compose — whichever is stricter wins.

Sub-agent delegation

Opt-in delegate_task hands a self-contained piece of work to an isolated sub-agent with its own context — never more permissive than the persona that spawned it.

AI PR review

Reviews a pull request file by file — never a truncated diff — then posts line-anchored comments with one-click suggestions via gh, after an explicit confirm. Deep mode adds a verification pass that drops false positives before you see them.

Web search

Off until you configure it. DuckDuckGo needs no key; Brave, Tavily and Serper take one, and a self-hosted SearXNG keeps queries off third-party infrastructure entirely. Results are quarantined as untrusted, same as any fetched page.

Runs on models that can’t tool-call

Plenty of local builds have no tool support in their chat template and reject the request outright. Heap Code notices, switches to a text-based tool protocol, and keeps going — so a model that would otherwise fail mid-task still finishes one.

Terminal CLI

npm i -g @heaplabs/heapcode-cli — the same agent, permissions, and checkpoints in your terminal, plus a headless -p mode with JSON output for CI.

Team bundles

Export or import a project's memory, instructions, and skills as a single file, so a team shares one setup.

Local audit dashboard

See exactly what ran, what was approved or denied, and what got reverted — computed entirely on your machine, nothing sent anywhere to produce it.