derive(prv[, out]) // 32-byte private key -> 32-byte public key
sign(msg, prv, pub[, out]) // 32-byte block hash -> 64-byte signature
verify(sig, msg, pub) // -> boolean
-verify_blocks(pub, blocks) // { hash, signature }[], up to 32 -> boolean[]
+verify_blocks(pub, blocks) // { hash, signature }[], up to 64 -> boolean[]
```
Inputs are `Uint8Array` or hex strings, and the output type follows the input
public key. It is correct across the full range and has vector coverage. The
public-key decompression is hoisted into `crypto_verify_decodepubkey` and run
once per batch, which saves `(1 - 1/n) * k` where `k` is the key work's share of
-one verification: **measured k = 8.84%**, realising **8.51%** at the cap of 32.
+one verification: **measured k = 8.84%**, realising **~8.7%** at the cap of 64.
⚠️ **`verify` and `verify_blocks` have different acceptance sets, on purpose.**
batch hoist possible. Call it cold and you verify against whichever key ran last,
with no error. `crypto_verify_strict` does its own key handling inline.
-**Do not raise the cap to chase throughput.** `(1 - 1/n)` is 96.9% at n=32 and
-99.7% at n=341 — the remaining 2.8% of an 8.5% prize is under a quarter of a
-percent of runtime, and n=341 blocks the thread for ~24 ms against ~2.2 ms. The
-cap is a latency decision. More headroom would have to come from attacking the
-per-block double scalar multiplication, which is a different algorithm.
+**The cap is a latency decision, not a throughput one.** `(1 - 1/n)` is 96.9% at
+n=32, **98.4% at n=64**, 99.7% at n=341 — past ~64 you buy tenths of a percent of
+an 8.8% prize with milliseconds of blocked thread (~4.4 ms at 64, ~24 ms at 341).
+More throughput would have to come from attacking the per-block double scalar
+multiplication, which is a different algorithm entirely.
**`sign` costs two fixed-base scalar multiplications, deliberately.** It derives
the public key from the private key and throws `Invalid public key` if the one
| `INPUT_SIG` | 64 B, one signature for single `verify` |
| `OUTPUT_DERIVE` | 32 B, a public key |
| `OUTPUT_SIGN` | 64 B, a signature |
-| `OUTPUT_VERIFY` | 32 B, one result byte per verified block (= `MAX_VERIFY_BLOCKS`) |
+| `OUTPUT_VERIFY` | 64 B, one result byte per verified block (= `MAX_VERIFY_BLOCKS`) |
Pointer getters are `ptrInputMsg`, `ptrInputPrv`, `ptrInputPub`, `ptrInputSig`,
`ptrOutputDerive`, `ptrOutputSign`, `ptrOutputVerify`.
Raising the *contract* rather than sizing the buffer is what makes this safe:
the host already validates message length against `MAX_MESSAGE_BYTELENGTH`, so
-that one guard now covers batch packing too. **If you change `MAX_VERIFY_BLOCKS`,
-check `memory.size()`** — the message limit follows it once a batch exceeds a
-page, and pages are the expensive unit here.
+that one guard now covers batch packing too. Raising `MAX_VERIFY_BLOCKS` is free
+until a batch exceeds a page — that is n = 682 — so the 32 → 64 bump moved only
+`OUTPUT_VERIFY`, by 32 bytes. Past 682 the message limit starts following the cap;
+**check `memory.size()` if you go there.**
These sizes are a **public API**: `nano25519.constants` carries
`BLOCKHASH_BYTELENGTH`, `KEY_BYTELENGTH`, `SIGNATURE_BYTELENGTH`,
`src/lib` and deliberately not exported from the package — the buffer-scrub
guarantee is a property of the module, not a promise about callers.
-`test/index.html` still hardcodes the cap as `32` in two places and is the one
-file known to violate this; `constants.MAX_VERIFY_BLOCKS` is now available to
-it.
+`test/index.html` reads `nano25519.constants.MAX_VERIFY_BLOCKS` for its loop
+bound, its per-block divisor and its default test size — which is why the 32 → 64
+cap change needed no edit there. Keep it that way: a stale denominator in the
+benchmark does not error, it silently reports a wrong per-block time in the file
+whose whole job is producing that number.
## Build constraints that break normal assumptions
- **`enable: ["simd"]`** — field elements are 12 `i32` limbs, not 10. The two
trailing limbs are padding; respect the stride.
- **`initialMemory: 4`** — pins the module at 4 pages (262,144 B); heap at rest
- is 135,824 B. Only the stub allocator's *growth* doubles (1→2→4→8); the
- setting itself takes any integer, so it need not be a power of two — **3 pages
- currently fit**, with 60,784 B spare. Re-derive this whenever a buffer is
- resized rather than leaving it stale or rounding it up; it has been wrong in
- both directions twice.
+ is 135,856 B. Only the stub allocator's *growth* doubles (1→2→4→8); the setting
+ itself takes any integer, so it need not be a power of two, and 3 pages do fit
+ today with 60,752 B spare. **Slack here is nearly free**: measured, the fourth
+ page changes neither throughput nor instantiation time and never becomes
+ resident, because nothing writes to it — it costs address space and V8
+ accounting, not RAM. Re-derive it whenever a buffer is resized rather than
+ leaving it stale, but note *stale* means smaller than the statics. Too large is
+ cheap; too small grows memory at init.
Memory growth *during* a call would detach every `Uint8Array` view the host
holds, so keeping call paths allocation-free is a correctness requirement, not