You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: content/notes/2026-09-16-typesafe-system-one-jev.md
+6-2Lines changed: 6 additions & 2 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -10,8 +10,8 @@ source_type = "article"
10
10
newsletter_candidate = true
11
11
why_it_matters = "TypeSafe is trying to make AI feel less like a chatbot and more like a typed, low-latency decision primitive that normal software can compose."
retrieval_note = "Launch article, docs introduction, and workflow evals page were extracted directly. Archer Hume's black-box reconstruction was added later and read directly."
15
15
+++
16
16
17
17
**Logged at IST:** 2026-09-16 07:45 IST
@@ -24,4 +24,8 @@ The docs make the software shape clearer: ask narrow atomic questions, evaluate
24
24
25
25
The interesting bet is architectural. If chat models are optimized for human-facing System 2-ish generation, TypeSafe is carving out System 1-ish model calls: low-latency, typed, probabilistic judgments embedded inside larger deterministic systems.
26
26
27
+
**Update, 2026-09-18:** Archer Hume's black-box reconstruction makes the architectural hypothesis much more concrete. The post argues Jev likely uses a causal transformer as a shared-state encoder, then evaluates isolated question branches against that shared state and reads typed probability distributions directly instead of decoding JSON text. The evidence is behavioural rather than a disclosure: token accounting appears additive, sibling questions cannot leak facts into one another while facts in shared state are visible, many questions stay cheap enough to suggest shared computation, and option-order / added-option probes show the alternatives are processed listwise rather than as independent fixed logits.
28
+
29
+
The post is careful about uncertainty: direct numerical readouts are supported by TypeSafe's own claims, question isolation and option interaction are observable behaviours, while KV sharing, pointer-style readouts, causal masking, and sparse MoE are progressively more speculative explanations. The useful takeaway is still strong: Jev is best understood as a decision-service architecture, not just a JSON classifier wrapper around a chat model.
30
+
27
31
**Newsletter angle:** A strong artifact for the "AI as software primitive" lane: intelligence as fast, typed decision nodes inside workflows, not only chat, copilot, or agent loops.
Copy file name to clipboardExpand all lines: content/notes/2026-09-18-hacktron-hacking-openai.md
+3-1Lines changed: 3 additions & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -11,7 +11,7 @@ newsletter_candidate = true
11
11
why_it_matters = "A concrete case study in how agent-assisted exploit work, a native image-decoder bug, and loose identity/connectors can turn into cross-product access quickly."
retrieval_note = "S1r1us and Harsh Jaiswal X posts were extracted via FXTwitter; attached images and related screenshots were OCR'd; Hacktron's OpenAI write-up, the HEIF Heist site, and Discourse GHSA-vhm9-85gw-x335 were read directly. A related LiveOverflow YouTube explainer was identified via oEmbed/X, but transcript extraction was blocked by YouTube IP restrictions, so the video itself was not summarized."
14
+
retrieval_note = "S1r1us and Harsh Jaiswal X posts were extracted via FXTwitter; attached images and related screenshots were OCR'd; Hacktron's OpenAI write-up, the HEIF Heist site, and Discourse GHSA-vhm9-85gw-x335 were read directly. The related LiveOverflow YouTube explainer transcript was extracted from YouTube's web UI after normal transcript APIs and mirror sites were blocked, empty, login-gated, or unavailable."
15
15
+++
16
16
17
17
**Logged at IST:** 2026-09-18 10:02 IST
@@ -26,6 +26,8 @@ The Discourse advisory gives the concrete patch hook: **GHSA-vhm9-85gw-x335 / CV
26
26
27
27
The broader HEIF Heist site frames the issue as an ecosystem problem rather than a one-off Discourse bug. Many products accept HEIF, HEIC, or AVIF uploads that eventually flow into native decoders such as `libheif` and `libde265` through ImageMagick, libvips, Sharp, distro packages, and container images. Hacktron lists related impact across OpenAI, Slack, Meta, Discourse, Next.js image optimization, GitHub Enterprise, and other frameworks/CMSes.
28
28
29
+
LiveOverflow's transcript adds two useful details. The original video apparently included a recording of the OpenAI internal PR proof, but LiveOverflow says OpenAI asked them not to show it, so the published video uses a reenactment. The technical arc matches the written sources: HEIF upload handling in Discourse routes through ImageMagick, `libheif` had an accidentally-fixed-but-untracked heap-overflow path, agent-assisted exploit adaptation bridged local Discourse to OpenAI's hosted environment, and the proof stopped at a harmless Codex PR rather than reading private repository contents.
30
+
29
31
The deeper point is economic: Hacktron argues that agent-assisted exploit development compressed work that used to require rare expertise and sustained effort into a few days of agent time plus a few hours of human guidance. That makes old assumptions about "known but hard to weaponize" dependency bugs much weaker.
30
32
31
33
**Newsletter angle:** Treat image decoding as an untrusted native execution boundary. Patch `libheif`/`libde265` and affected applications, but also sandbox or disable untrusted HEIF/AVIF processing where it is not needed, especially when the surrounding product has powerful identity and connector reach.
why_it_matters = "If the reported retention holds up, 27B-class reasoning and multimodal models are getting small enough for local assistants, private workflows, and single-GPU serving without a large quality cliff."
retrieval_note = "PrismML X launch post was extracted via FXTwitter; the attached benchmark image was OCR'd; PrismML's launch post, Hugging Face model card, and TechCrunch coverage were read directly."
15
+
+++
16
+
17
+
**Logged at IST:** 2026-09-18 10:58 IST
18
+
19
+
**What it is:** PrismML's launch of Ternary Bonsai 2 27B, a compressed Qwen3.8-27B-derived multimodal model released under Apache 2.0.
20
+
21
+
**Gist:** PrismML says Bonsai 2 27B uses ternary `{−1, 0, +1}` weights with FP16 group-wise scaling, landing around **1.76 effective bits per weight** and a **5.9 GB** footprint. The company claims this is more than **9x smaller** than the full-precision Qwen3.8 27B counterpart while retaining **98.2%** of aggregate benchmark performance.
22
+
23
+
The headline table is useful because it shows where the loss remains. Overall score is **83.9** versus Qwen3.8 27B at **85.4**; math is nearly preserved (**96.57** versus **97.06**), coding is close (**81.58** versus **82.17** in the launch table), agentic/tool calling drops more (**77.57** versus **79.74**), and vision drops from **81.64** to **78.59**. Instruction-following is reported as slightly above the base model at **82.66** versus **81.25**.
24
+
25
+
The Hugging Face card adds the deployment shape: GGUF packs for PrismML's `llama.cpp` fork, MLX companion weights for Apple Silicon, 262K context inherited from the Qwen3.8-27B hybrid-attention backbone, and benchmark comparisons against conventional low-bit quantizations. PrismML frames the result as an intelligence-density gain rather than only a size reduction: a 27B-class model that fits on a normal laptop or a single consumer/datacenter GPU.
26
+
27
+
**Newsletter angle:** Local AI is moving from "small model because hardware is constrained" toward "large model behaviour under a small memory and power budget." The interesting metric is becoming useful capability per GB / joule, not only benchmark score at full precision.
0 commit comments