ResearchOS/Wiki

Security

Your notebook is stored in a folder you control. Account services and optional sharing, AI, identity-key backup, and publishing have their own data flows. This page explains the distinction and how to inspect network activity.

The claim

ResearchOS stores your local notebook in a folder you choose. It also has servers for sign-in, accounts, billing, sharing, live collaboration, companion devices, and hosted sites. Using a local notebook does not automatically upload its contents to those services. Features you enable can send selected content, and a cloud provider you use to sync the folder can store a copy. The privacy policy describes those services and the data they process.

What stays on your computer

When you pick a data folder, ResearchOS reads and writes your notebook directly to that folder on your machine. Optional online features have separate data flows. You can open the folder at any time in your operating system's file browser and see exactly what is there.

Here is what the folder holds.

  • Every experiment, lab note, and result you write.
  • Every image, PDF, and arbitrary attachment you drop into a note.
  • The JSON for projects, tasks, dependencies, methods, and PCR protocols.
  • Your calendar subscription URLs and the events they pulled in.
  • The Argon2id-protected credentials in _account.json guarding your per-user profile.
Show me the exact file layout
<your data folder>/
  .gitignore                     app-managed, auto-appended for sensitive sidecars
  users/
    _user_metadata.json          cross-user color preferences + display names
    _global_counters.json        cross-user id allocator
    public/                      methods + PCR protocols shared across all users
    lab/                         shared lab-account state (funding accounts, etc.)
    <username>/
      settings.json              user settings, default views, tab visibility
      _account.json              Argon2id-protected credentials (replaces _auth.json)
      _counters.json             id counters per entity type
      _shared_with_me.json       inbound sharing pointers
      _calendar-feeds.json       subscribed iCal URLs
      _notifications.json        bell-dropdown rows
      _history/                  per-record version history (.jsonl append logs)
        task/<id>.jsonl
        task_notes/<id>.jsonl
        task_results/<id>.jsonl
        project/<id>.jsonl
        notes/<id>.jsonl
        sequences/<id>.jsonl
        molecules/<id>.jsonl
        inventory_items/<id>.jsonl
      projects/<id>.json         one flat file per project
      tasks/<id>.json            one flat file per task (the JSON carries
                                 fields; long-form text and attachments
                                 live under results/task-<id>/ below)
      methods/<id>.json          reusable protocols
      pcr_protocols/<id>.json    PCR programs and recipes
      lc_gradients/<id>.json     LC gradient programs
      cell_culture_schedules/<id>.json
      plate_layouts/<id>.json
      notes/<id>.json            shared lab notes
      sequences/<id>.json        DNA/RNA/protein sequences
      molecules/<id>.json        chemical structures
      inventory_items/<id>.json  inventory catalog entries
      datahub/<id>.json          Data Hub datasets
      phylo/<id>.json            phylogenetic tree files
      figures/<id>.json          figure composer artboards
      results/task-<id>/         per-task long-form content + attachments
        notes.md                 your Lab Notes tab writeup
        results.md               your Results tab writeup
        notes/Images/            images dropped into Lab Notes tab
        notes/Files/             files dropped into Lab Notes tab
        results/Images/          images dropped into Results tab
        results/Files/           files dropped into Results tab
      events/<id>.json           native calendar events
      goals/<id>.json            project goals
      lab_links/<id>.json        link library entries
      purchase_items/<id>.json   purchase orders
      inbox/Images/              photos awaiting routing

Online services and network requests

The following examples are not an exhaustive list of backend routes. Account and billing services retain service records. Sharing and device relays handle content according to the feature, and publishing creates a public hosted copy. AI requests send the conversation and selected context to the configured inference provider. See the privacy policy before sending sensitive content through an online feature.

  • Calendar feed sync. When you subscribe to an iCal URL (Google, Outlook, iCloud, a university calendar), the browser asks /api/calendar-feed to fetch the iCal text from the upstream and stream it back. The subscription URL travels in the x-calendar-url request header rather than in the URL query string, so it is never written to Vercel access logs. A 15-minute edge cache keeps repeated polls from hammering the upstream. We do not persist the URL or the contents.
  • ORCID publication lookup. When you view a researcher profile that lists an ORCID iD (during sharing or profile setup), the browser asks /api/orcid/works to fetch that person's public publication list from the ORCID API and hand it back. ORCID's API isn't CORS-open, so the call has to go server-side. The route only ever touches public ORCID data, it's rate-limited per IP, and it stores nothing.
  • BeakerBot AI assistant. When you send a message to the built-in AI assistant (BeakerBot), the browser calls /api/ai/chat, which forwards the conversation to the configured inference provider and streams the response back. This is optional. Requests can contain your messages and local records selected as context by you or the assistant tools. Usage metering records token consumption separately from your local conversation history. Review the provider and privacy settings before sending research content.
  • Vercel Web Analytics. When you navigate between pages, your browser sends an anonymous page-view beacon to Vercel telling them which route you visited. No IDs, no folder contents, no typed text, no markdown bodies, no project names. Vercel sees your IP address (which they hash before storage per their privacy policy). The Settings → Offline mode toggle disables it durably, the script tag is never injected when offline mode is on.
  • Feature-usage beacon. When you use a tracked feature like sending a share or publishing a directory profile, the browser fires a small anonymous beacon to /api/analytics/event so the operator dashboard can count how often features get used. The payload is allow-listed enum and boolean flags only (was an ORCID present, did the share go to an existing user), never an ID, name, title, or anything you typed. It rides the same Offline mode gate, and the server re-validates and stores only the allow-listed shape.

The calendar CORS-bypass proxy uses the most defensive shape we know how to write, with HTTPS only, private-IP blocking, redirect re-validation, byte cap, timeout, content-type denylist, and per-IP rate limiting. The Vercel Analytics endpoint is a Vercel-owned script and beacon target, its posture is Vercel's, not ours. The route-defense code is in frontend/src/lib/api/url-guards.ts and frontend/src/lib/api/rate-limit.ts if you want to read it line by line.

What we collect, and what we don't

We collect anonymous page-view pings via Vercel Web Analytics. When you navigate between pages, anonymous beacons go to Vercel. No IDs, no folder contents, no typed text, no markdown bodies. We use this to know which pages get used and which sit idle. Settings → Offline mode disables it durably, the script is not injected when the toggle is on, and the toggle is read at component-mount time so the choice survives reloads.

We also collect anonymous feature-usage counts via /api/analytics/event, such as how often a share gets sent. These carry allow-listed enum and boolean flags only, never an ID, name, or anything you typed, and they ride the same Offline mode gate.

We do not collect anything else. No Sentry, no Google Analytics, no Mixpanel, no PostHog, no Hotjar, no Datadog, no Amplitude. No background phone-home. No crash reporter. No content telemetry. Running npm ls against the repo will confirm only @vercel/analytics is present, and the network tab will confirm no other endpoints are contacted.

The Report an issue button does not auto-submit anything. When you click it, your browser opens a pre-filled GitHub issue URL in a new tab. You see the body, edit it, and click Submit. Nothing happens until you do.

The Report an Issue modal. You see the body, you edit the body, you choose whether to submit.

Honest limits worth knowing about

These claims hold, but they have implications worth understanding before you trust ResearchOS with anything sensitive.

How to verify it yourself

You don't have to take any of this on faith. There are two ways to confirm what the app is doing. The first is built into Settings and takes about thirty seconds.

Data inventory panel in Settings. Every file path and every outbound endpoint is listed here.
Offline mode toggle in Settings. One click disables the analytics script and all proxy routes.
The thorough way, open DevTools and watch the network yourself

Your browser already shows every network request the app makes. This is the audit-grade path for anyone who wants to see the bytes themselves.

  1. Open ResearchOS in Chrome or Edge, then open DevTools. F12 works on Windows and Linux. Cmd + Option + I works on macOS.
  2. Switch to the Network tab and check Preserve log so refreshes don't clear the view.
  3. Reload the page. Visit Calendar, the inbox, and the experiments you care about. Watch every outbound request as you go.
  4. You should see requests to your own ResearchOS origin (for JavaScript, CSS, and static assets), occasional requests to /api/calendar-feed when a subscribed feed refreshes, a one-off request to /api/orcid/works if you open a researcher profile that lists an ORCID iD, and occasional requests to va.vercel-scripts.com and vitals.vercel-insights.com for anonymous page-view pings, plus an occasional anonymous /api/analytics/event beacon when you use a tracked feature like sending a share (allow-listed enum and boolean props only, no IDs or text). All of those stop when Offline mode is on. One more destination may appear in a narrow circumstance. If the AI Helper prompts bundled with your running app are older than the latest deploy and you click Pull latest from research-os-xi.vercel.app in Settings → AI Helper, the browser fetches https://research-os-xi.vercel.app/ai-helper/manifest.json and https://research-os-xi.vercel.app/ai-helper/{size}.md. This is a user-initiated, on-demand pull, not a background call. Nothing else.
  5. For a second pass, switch from the Network tab to Application → IndexedDB. ResearchOS uses two IndexedDB databases.
    • research-os-fsa holds the opaque File System Access handle for your data folder. This is what gives the browser permission to read and write the folder you picked.
    • keyval-store holds three small session-routing strings, the folder name plus its grant timestamp, the currently signed-in user, and (if you are a PI signed in) the primary account.

What we just fixed

ResearchOS went through an internal security audit on 2026-05-15. It closed one Critical finding (cross-site scripting via unsanitized markdown HTML across all 8 markdown rendering sites in the app) and several Important findings around the proxy routes, the LabArchives credential flow, and the in-app verification surface (data inventory, offline mode, storage hints). The audit report is checked into the repo as SECURITY_AUDIT.md, and the merge batch on main sits around 94f0ab08 (audit doc) and 813748a5 (XSS sanitize + CSP).

If you find something that looks wrong, open a GitHub issue (the Report an issue button is the fastest path) or email the maintainer directly. We'd much rather hear about a problem early.

Atomic-write safety

A subtler kind of data loss has nothing to do with the network and everything to do with timing. If an app is halfway through writing a file when the tab closes, the laptop sleeps, or the process crashes, a naive implementation can leave the file zero-byte or truncated, and the good version you had a second ago is gone. ResearchOS is built so that cannot happen.

Every save writes to a temporary .tmp file first, lets that file finish landing durably on disk, and only then atomically moves it into place over the real file. The move is the kind of operation that either fully happens or does not happen at all, with no in-between state on disk. So if a crash interrupts a save, the worst case is that your previous good version survives untouched. You can never be left with a half-written or empty file in place of your data. The implementation is in frontend/src/lib/file-system/file-service.ts if you want to read the exact sequence.

Tested on every commit

A trust claim is only as good as the thing that keeps it true over time. The whole app is gated by automated tests that run on every commit and every pull request to main. That gate runs linting, a full TypeScript typecheck, the Vitest unit and integration suite with coverage, and Playwright end-to-end tests in a real browser. The workflow lives in .github/workflows/ci.yml. If any of those gates fail, the change does not ship. That same machinery is what keeps the scientific calculations honest, which is its own page. Method validation explains how every sequence and lab calculation is re-checked against the reference tools the field trusts, on every commit.