Repository navigation
[codex] test Agent OS mountFs layers - #1521
Merged
Merged
Conversation
railway-app
Bot
temporarily deployed
to
rivet-frontend / agentos-pr-1521
June 24, 2026 23:15
Destroyed
|
🚅 Deployed to the agentos-pr-1521 environment in agentos
🚅 Deployed to the agentos-pr-1521 environment in rivet-frontend
|
railway-app
Bot
temporarily deployed
to
rivet-frontend / agentos-pr-1521
June 25, 2026 00:28
Destroyed
railway-app
Bot
temporarily deployed
to
rivet-frontend / agentos-pr-1521
June 25, 2026 01:41
Destroyed
railway-app
Bot
temporarily deployed
to
rivet-frontend / agentos-pr-1521
June 25, 2026 01:52
Destroyed
Runtime mountFs/unmountFs each trigger a full configureVm reconfigure; the two tests that do both sequentially exceed vitest's 30s default in CI (observed 30-86s). Give them an explicit 120s budget so CI stops flaking. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
railway-app
Bot
temporarily deployed
to
rivet-frontend / agentos-pr-1521
June 25, 2026 07:23
Destroyed
NathanFlurry
marked this pull request as ready for review
June 25, 2026 07:51
NathanFlurry
added a commit
that referenced
this pull request
Jun 25, 2026
…1524) * Add Agent OS homepage diagrams: harness architecture + cold-start comparison Two native React + inline SVG + Tailwind + framer-motion visuals on the Agent OS homepage, themed to the site palette: - HarnessArchitecture: a cross diagram framed as Agent OS — a cycling agent logo at the center routes requests/responses to Tools, Session, Sandbox, and Orchestration with animated flow dots; collapses to a stacked layout on mobile. - ColdStartRace: a containers-vs-Agent OS cold-start comparison. Each container boots on its own (red→green + mini bar) carrying its own ~1 GB; Agent OS packs all agents into one shared process at ~131 MB each. Includes live ms counters, a speed slider to control playback, and a stats row (92x faster, 8x less memory, 32x cheaper) sourced from bench.ts. Also bundles pre-existing in-progress site changes from this workspace: Navigation, Footer, registry (RegistryPageClient, registry.ts, registry-icons.ts, registry/[slug].astro), pricing page, and favicon. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Homepage: animated performance benchmarks + architecture copy fixes Replace the static benchmark cards with three animated, light-styled diagrams (cold-start race with p50/p95/p99 toggle, memory-overhead squares, execution-density packing), extract shared benchUI primitives (BenchToggle/CountUpStat/BenchInfoTooltip), add per-tier execs/costPerHour to bench.ts, link hero stats to their diagrams, and fix the info tooltips and the agentOS boot timing. Correct the architecture framing to match the product: an OS for all AI agents and arbitrary code, many lightweight in-process VMs packed into one process on an in-process kernel. Also bundles related in-progress homepage edits (navigation, footer, registry, pricing, use-cases, config). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Homepage: rework feature sections, real syntax highlighting, remove pricing - Reorganize the two overlapping homepage feature sections. "Meet your agent's new operating system" now maps 1:1 to the harness architecture diagram (Agent + Tools/MCP, Session, Sandbox, Orchestration), and the "Everything agents need, built in" bento covers distinct cross-cutting capabilities (fault isolation, CPU/memory limits, network control, filesystem mounts, embed-in-backend, deploy-anywhere) with no duplication. - Replace the plain code emitter with a TS/JS tokenizer so the hero code tabs get real inline-colored syntax highlighting (highlight-code.ts). - Drop monospace from UI labels: `font-mono` now renders the sans stack; real code/terminal blocks use the new `font-code` (JetBrains Mono). - Remove the pricing page and its nav/footer links (AgentOSPricingPage, pricing.astro, agent-os-pricing faqs). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Homepage hero: inline benchmark highlights + monospace code blocks Replace the separate benchmark stat cards with compact inline highlights under the tagline (linked to the benchmarks below), move "Works with" in line with the CTA buttons, and trim now-unused imports/state. Force the shiki syntax-highlighter output back to a real monospace font, since the Tailwind `font-mono`→Manrope mapping was leaking into code blocks. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Homepage: mobile harness diagram — hub-and-spoke to match desktop The mobile layout stacked the cards in a linear chain (Agent → Tools → Session → Sandbox → Orchestration), reading as a sequential pipeline and contradicting both the desktop cross and the diagram's own aria-label. Redesign it as a "comb": a single trunk descends from the Agent hub with four horizontal branches, each carrying a request (solid accent) and response (dashed ink) arrow, so every service routes through the agent. Spacing between the hub and the first branch matches the rest of the stack. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Homepage harness card: larger, balanced floating agent logos Scale up the floating agent logos in the "Run any agent harness" bento card (max tile 62→84px, nothing under 64px) and redistribute them into a balanced arc that sweeps top-right → down the right edge → across the bottom, weighting the cluster right-of-center to counter the left-aligned copy. Desktop-only logos; no text overlap down to the lg breakpoint. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Homepage: move hero code into a dedicated orchestration section Rework the hero into a two-column layout — brand, value prop, benchmark highlights, and CTAs on the left; a live agent-session terminal on the right — and split the old single tagline into a headline plus subhead. Move the tabbed code block out of the hero into a new AgentOrchestration section directly below, framed around orchestration (durable sessions, agent-to-agent, workflows, cron) with per-tab docs links. Reorder the hero tabs orchestration-first. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Homepage: redesign hero, code showcase, and OS/orchestration sections - Center the hero on a single column and remove the boot terminal - Rename CTA to "Set up with your agent"; "isolated Linux VM"; "agentOS runs" - Add a scroll-revealed "Orchestrate fleets…" title above the code showcase - Reorder sections: A new architecture → Operating system → Orchestration → Registry - Operating system: bento with floating agent logos on the "Any agent harness" tile - Orchestration: Multiplayer-led bento under the code showcase - Architecture copy: cite Google Chrome and Cloudflare Workers - Reveal the nav logo on scroll once the hero logo passes behind the nav Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * [codex] test Agent OS mountFs layers (#1521) * test agentos mount fs layers * docs: clarify filesystem mount support * fix core js vfs mountfs bridge * test and fix core api surface gaps * test: raise timeout for runtime mount/unmount integration tests Runtime mountFs/unmountFs each trigger a full configureVm reconfigure; the two tests that do both sequentially exceed vitest's 30s default in CI (observed 30-86s). Give them an explicit 120s budget so CI stops flaking. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix agentos rivetkit runtime issues (#1529) * perf(2a): pi SDK snapshot bundle (one snapshottable IIFE) esbuild bundles the lean pi-SDK graph (the 13 exports loadPiSdkRuntime needs) into a single IIFE, dist/pi-sdk-snapshot.js, that publishes globalThis.__PI_SDK_RUNTIME__. - src/snapshot-entry.ts: imports the SDK exports, assigns the runtime global. - scripts/build-snapshot-bundle.mjs: format:iife, platform:node (node builtins external), deep-import plugin to bypass the package exports map (avoids the ~100MB TUI graph the adapter skips), provider SDKs kept lazy/external, sha256 sidecar for dep-keying. C0 mitigations baked in: import.meta.url pinned to projected guest path; clipboard native-addon require made unreachable (DISPLAY=''). - package.json build chains build:snapshot; tsconfig excludes the esbuild-only entry. Validated (host, CJS harness): 13/13 exports populate, ~201ms eval, 880 modules inlined, 7.6MB. Traced top-level fs footprint = 2 deterministic package.json reads only (no net/ sockets/timers/native addons/per-session config). Artifact (dist) is gitignored. Step 2a of the agent-SDK snapshot goal. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * perf(2c+safety): pi-tui Intl.Segmenter Foreign fix + snapshot opt-in flag + docs Foreign fix (the last 2d blocker): build-snapshot-bundle.mjs snapshotSafePlugin now also lazy-inits pi-tui/dist/utils.js's module-level `new Intl.Segmenter` (an ICU-backed native External that V8 SnapshotCreator can't serialize) behind a Proxy — created on first use post-restore, not at module-init. With this, the real full pi SDK bundle snapshots cleanly and restores with createAgentSession/createAllTools intact (secure-exec Part 22 = ok). Also adds: config.js package.json inline, env-api-keys sync-require, and the PI_SNAPSHOT_ENTRY/OUTFILE/STUB_MODULES bisection overrides used to find it. Flag: agent.snapshot?: boolean on AgentSoftwareDescriptor (default false, opt-in, automatic fallback to per-session dynamic-import for SDKs that aren't snapshot-safe or fail to snapshot). Docs: agent.snapshot field + 'SDK snapshotting & snapshot-safety' rules in custom-software/definition.mdx. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * perf: agent SDK snapshot wiring + bench harness + preview secure-exec dep - Pin @secure-exec/* to the secure-exec preview (npm); @agentos-software/* preserved. - Clone-at-pinned-version cargo build mechanism (prepare-build) — crates.io has no preview track; clone secure-exec at the pinned sha and build cargo local for previews. Wired into publish.yaml + darwin.Dockerfile; documented in CLAUDE.md. - Rust client: add snapshot_userland_code: None to JsRuntimeConfig initializer. - Bundle transform hardening: env-api-keys transform asserts exactly 3 eager imports (loud fail on drift); pin @mariozechner/* exactly (direct + overrides). - JS-layer session benchmark + baseline gate + vm-vs-node phase traces. - Agent SDK snapshot opt-in (pi), MinimalResourceLoader fast path, architecture doc. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(CLAUDE): add Limits, Bounds & Observability standard (#1519) Codify the standard for agent-os-owned bounds (ACP/session/frame timeouts, event buffer, session/shell-id retention, host-tool registration caps, in-VM adapter log buffers): bounded-by-default, forward (don't reimplement) secure-exec limits, warn-on-approach + typed error for agent-os-owned bounds, no unbounded per-entity collections, and host-visible warnings via the agent log channel. Companion to the secure-exec CLAUDE.md standard. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(codex+claude): codex e2e agent + claude patch-minimization, re-fit onto main's @agentos-software/* layout (#1528) * perf(bench): make session.bench actually measure the agent-SDK snapshot The session benchmark was measuring the non-snapshot path on both axes, so it showed no difference between snapshot on/off: - loadPiSoftware() preferred the published @agentos-software/pi (no agent.snapshot, no dist/sdk-snapshot.js); now prefer the in-repo registry build (the code under test, which carries the snapshot bundle), falling back to the published package. - vm lane created a fresh sidecar per session, so the process-wide snapshot cache (build-once-per-sidecar) was never reused. Add a reused-sidecar lane (default) that creates one VM and loops createSession, with awaited session teardown so adapter processes don't pile up and confound the measurement. --fresh-vm-per-session keeps the old cold path. With a snapshot-capable sidecar + local pi, this measures createSession p50 ~405ms (snapshot on) vs ~994ms (off) in release. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(core): share one sidecar process per AgentOsSidecar handle (#1531) VMs leased from an AgentOsSidecar handle now run as incremental tenants of a single native sidecar process instead of spawning one process each. The handle owns the process (spawn-once, lazily); each VM's kernel proxy scopes its event pump to its own ownership and tears down only its VM on dispose (ownsClient), while the shared process is disposed with the handle. AgentOs.create() with no sidecar option uses the shared default pool, so this is the default everywhere (incl. RivetKit). Effect: per-VM memory drops to the marginal cost (~21MB shell vs a full per-VM process) and warm cold start drops to single-digit ms. Also: bench harness creates the sidecar up front + a cold-run snapshot, runs against the release sidecar, and uses statistically meaningful sample sizes; benchmarks.mdx documents how to run/reproduce. Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com> Co-authored-by: Nathan Flurry <NathanFlurry@users.noreply.github.com> Co-authored-by: Nathan Flurry <git@nathanflurry.com>
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Why
This verifies the mountFs/custom JS VFS path at the Agent OS core layer and documents the Rivet native plugin boundary, where callback-backed JS VFS drivers cannot cross the NAPI config envelope.
Companion secure-exec PR: rivet-dev/dynamic-apps#125
Validation