From bee3b719ba7b9eb3ef12781995351ca81a3579c0 Mon Sep 17 00:00:00 2001 From: Chris Duncan Date: Fri, 28 Aug 2026 14:56:45 -0700 Subject: [PATCH] Update agent file. --- AGENTS.md | 44 +++++++++++++++++++++++++------------------- 1 file changed, 25 insertions(+), 19 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index e9a34f3..185c8a9 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -14,7 +14,7 @@ boundary. Everything is synchronous and single-threaded. derive(prv[, out]) // 32-byte private key -> 32-byte public key sign(msg, prv, pub[, out]) // 32-byte block hash -> 64-byte signature verify(sig, msg, pub) // -> boolean -verify_blocks(pub, blocks) // { hash, signature }[], up to 32 -> boolean[] +verify_blocks(pub, blocks) // { hash, signature }[], up to 64 -> boolean[] ``` Inputs are `Uint8Array` or hex strings, and the output type follows the input @@ -25,7 +25,7 @@ trailing argument. public key. It is correct across the full range and has vector coverage. The public-key decompression is hoisted into `crypto_verify_decodepubkey` and run once per batch, which saves `(1 - 1/n) * k` where `k` is the key work's share of -one verification: **measured k = 8.84%**, realising **8.51%** at the cap of 32. +one verification: **measured k = 8.84%**, realising **~8.7%** at the cap of 64. ⚠️ **`verify` and `verify_blocks` have different acceptance sets, on purpose.** @@ -48,11 +48,11 @@ it reuses the module-level `A` that call leaves behind, which is what makes the batch hoist possible. Call it cold and you verify against whichever key ran last, with no error. `crypto_verify_strict` does its own key handling inline. -**Do not raise the cap to chase throughput.** `(1 - 1/n)` is 96.9% at n=32 and -99.7% at n=341 — the remaining 2.8% of an 8.5% prize is under a quarter of a -percent of runtime, and n=341 blocks the thread for ~24 ms against ~2.2 ms. The -cap is a latency decision. More headroom would have to come from attacking the -per-block double scalar multiplication, which is a different algorithm. +**The cap is a latency decision, not a throughput one.** `(1 - 1/n)` is 96.9% at +n=32, **98.4% at n=64**, 99.7% at n=341 — past ~64 you buy tenths of a percent of +an 8.8% prize with milliseconds of blocked thread (~4.4 ms at 64, ~24 ms at 341). +More throughput would have to come from attacking the per-block double scalar +multiplication, which is a different algorithm entirely. **`sign` costs two fixed-base scalar multiplications, deliberately.** It derives the public key from the private key and throws `Invalid public key` if the one @@ -105,7 +105,7 @@ every pointer and byte length once at module load and re-exports them. | `INPUT_SIG` | 64 B, one signature for single `verify` | | `OUTPUT_DERIVE` | 32 B, a public key | | `OUTPUT_SIGN` | 64 B, a signature | -| `OUTPUT_VERIFY` | 32 B, one result byte per verified block (= `MAX_VERIFY_BLOCKS`) | +| `OUTPUT_VERIFY` | 64 B, one result byte per verified block (= `MAX_VERIFY_BLOCKS`) | Pointer getters are `ptrInputMsg`, `ptrInputPrv`, `ptrInputPub`, `ptrInputSig`, `ptrOutputDerive`, `ptrOutputSign`, `ptrOutputVerify`. @@ -139,9 +139,10 @@ export const MAX_MESSAGE_BYTELENGTH: i32 = Raising the *contract* rather than sizing the buffer is what makes this safe: the host already validates message length against `MAX_MESSAGE_BYTELENGTH`, so -that one guard now covers batch packing too. **If you change `MAX_VERIFY_BLOCKS`, -check `memory.size()`** — the message limit follows it once a batch exceeds a -page, and pages are the expensive unit here. +that one guard now covers batch packing too. Raising `MAX_VERIFY_BLOCKS` is free +until a batch exceeds a page — that is n = 682 — so the 32 → 64 bump moved only +`OUTPUT_VERIFY`, by 32 bytes. Past 682 the message limit starts following the cap; +**check `memory.size()` if you go there.** These sizes are a **public API**: `nano25519.constants` carries `BLOCKHASH_BYTELENGTH`, `KEY_BYTELENGTH`, `SIGNATURE_BYTELENGTH`, @@ -151,9 +152,11 @@ boundary and in consumers. Buffer **pointers** are grouped as `pointers` inside `src/lib` and deliberately not exported from the package — the buffer-scrub guarantee is a property of the module, not a promise about callers. -`test/index.html` still hardcodes the cap as `32` in two places and is the one -file known to violate this; `constants.MAX_VERIFY_BLOCKS` is now available to -it. +`test/index.html` reads `nano25519.constants.MAX_VERIFY_BLOCKS` for its loop +bound, its per-block divisor and its default test size — which is why the 32 → 64 +cap change needed no edit there. Keep it that way: a stale denominator in the +benchmark does not error, it silently reports a wrong per-block time in the file +whose whole job is producing that number. ## Build constraints that break normal assumptions @@ -176,11 +179,14 @@ ordinary AssemblyScript: - **`enable: ["simd"]`** — field elements are 12 `i32` limbs, not 10. The two trailing limbs are padding; respect the stride. - **`initialMemory: 4`** — pins the module at 4 pages (262,144 B); heap at rest - is 135,824 B. Only the stub allocator's *growth* doubles (1→2→4→8); the - setting itself takes any integer, so it need not be a power of two — **3 pages - currently fit**, with 60,784 B spare. Re-derive this whenever a buffer is - resized rather than leaving it stale or rounding it up; it has been wrong in - both directions twice. + is 135,856 B. Only the stub allocator's *growth* doubles (1→2→4→8); the setting + itself takes any integer, so it need not be a power of two, and 3 pages do fit + today with 60,752 B spare. **Slack here is nearly free**: measured, the fourth + page changes neither throughput nor instantiation time and never becomes + resident, because nothing writes to it — it costs address space and V8 + accounting, not RAM. Re-derive it whenever a buffer is resized rather than + leaving it stale, but note *stale* means smaller than the statics. Too large is + cheap; too small grows memory at init. Memory growth *during* a call would detach every `Uint8Array` view the host holds, so keeping call paths allocation-free is a correctness requirement, not -- 2.52.0