node-re2 provides RE2 regular expression bindings for Node.js. Prior to version 1.26.1, passing a Buffer whose final bytes form a truncated…
[email protected]·CWE-125·Published 2026-08-06
node-re2 provides RE2 regular expression bindings for Node.js. Prior to version 1.26.1, passing a Buffer whose final bytes form a truncated (incomplete) multi-byte UTF-8 sequence could cause the native binding to read past the end of the allocated buffer while attempting to decode the final, incomplete code point. This could result in an out-of-bounds read and potential disclosure of adjacent memory contents. This issue is fixed in version 1.26.1.
## Summary `re2` infers a character's byte length from its UTF-8 lead byte alone, with no bound on the bytes actually remaining in the input. `Buffer` arguments reach the native layer verbatim — only strings are re-encoded into well-formed UTF-8 — so a `Buffer` whose last byte is a multi-byte lead promises continuation bytes that are not there, and the result builders read up to 3 bytes past the end of the buffer. In `replace()` and `split()` those bytes are copied into the returned `Buffer`, disclosing adjacent heap memory to JavaScript. The trigger is deterministic and requires no special heap grooming. Only `Buffer` input is affected. String input was never at risk: re-encoding guarantees every multi-byte sequence is complete. ## Root cause `getUtf8CharSize` maps a lead byte to a length of 1–4 and never sees the input size: ```cpp // lib/wrapped_re2.h inline size_t getUtf8CharSize(char ch) { return ((0xE5000000 >> ((ch >> 3) & 0x1E)) & 3) + 1; } ``` Callers then read that many bytes. In the zero-width branch of `replace()`, the guard proves only that at least *one* byte remains: ```cpp // lib/replace.cc else if ((size_t)offset < size) { auto sym_size = getUtf8CharSize(data[offset]); // may claim up to 4 bytes result.append(data + offset, sym_size); // reads data[offset .. offset + 3] byteIndex = offset + sym_size; } ``` `offset < size` permits `offset == size - 1`, so a lead byte of `0xF0` makes `append` read `data[size]`, `data[size + 1]` and `data[size + 2]`. Seven read sites shared the defect: | Site | Argument | Disclosed to JS | |---|---|---| | `lib/replace.cc` (zero-width branch) | subject | yes | | `lib/replace.cc` (callback replacer) | subject | yes | | `lib/replace.cc` (replacement scan) | replacement | yes | | `lib/split.cc` | subject | yes | | `lib/pattern.cc` `translateRegExp` (x2) | pattern | no | | `lib/pattern.cc` `escapeRegExp` | pattern | no | Three further callers were **not** vulnerable, because they use the result only to advance an index and never dereference past the end: `getUtf16PositionByCounter` in `lib/wrapped_re2.h` (clamps its return to the buffer size), `lib/match.cc` (the value feeds `RE2::Match`, which rejects `startpos > endpos`), and the `getMaxSubmatch` scan in `lib/replace.cc` (an overshoot just ends the loop). ## Proof of concept Each call returns more bytes than were supplied; the trailing bytes are heap contents and vary between runs. ```js const RE2 = require('re2'); const hex = buf => [...buf].map(b => b.toString(16).padStart(2, '0')).join(' '); // subject: 2 bytes in, 5 bytes out console.log(hex(new RE2('', 'g').replace(Buffer.from([0x41, 0xf0]), ''))); // 41 f0 61 7b eb <- last 3 bytes are adjacent heap memory // replacement argument console.log(hex(new RE2('A', 'g').replace(Buffer.from('A'), Buffer.from([0x42, 0xf0])))); // 42 f0 41 26 d6 // split console.log(new RE2('', 'g').split(Buffer.from([0x41, 0xf0])).map(hex)); // [ '41', 'f0 e2 e4 df' ] ``` `0xC2` (2-byte lead) and `0xE2` (3-byte lead) over-read 1 and 2 bytes respectively; `0xF0` over-reads 3. For the pattern path the over-read occurs in `translateRegExp` / `escapeRegExp`, which run before RE2 validates the pattern, but RE2 then rejects the malformed input, so the bytes are discarded rather than returned: ```js new RE2(Buffer.from([0xf0])); // SyntaxError: invalid UTF-8 — read already happened ``` ## Impact **Information disclosure (`replace`, `split`).** Up to 3 bytes of heap memory adjacent to the input buffer are returned to JavaScript per call. The read is repeatable, so an attacker who controls `Buffer` input and observes output can sample heap memory incrementally. What lands there depends on allocator layout and is not directly steerable, but it may include fragments of other buffers. **Out-of-bounds read (pattern compilation).** No disclosure path, since the malformed pattern is rejected — but the read is still undefined behavior and can fault if the buffer ends on a page boundary. Applications that pass only strings, or only well-formed UTF-8 buffers, are unaffected. The exposure matters most where `re2` is used as intended: running patterns or subjects derived from untrusted input. ## Suggested fix Clamp the inferred character size to the bytes that actually remain, at every site whose result indexes the buffer: ```cpp inline size_t getUtf8CharSize(char ch, size_t remaining) { size_t size = getUtf8CharSize(ch); return size < remaining ? size : remaining; } ``` This is O(1) and changes no algorithm's complexity. A truncated tail then round-trips as the bytes it really holds, which preserves the documented contract that `Buffer` input is passed through verbatim. Rejecting malformed UTF-8 in `Buffer` input would also close the hole, but is a breaking API change. ## Resolution Fixed in `[email protected]`. All seven read sites now clamp the character size to the remaining input, so a `Buffer` ending in a truncated multi-byte character round-trips as its own bytes instead of reading past the end. Regression tests cover the subject, replacement and pattern positions for 2-, 3- and 4-byte leads, including partially truncated sequences. **Remediation:** upgrade to `[email protected]` or later. **Workaround** (if you cannot upgrade): pass strings rather than `Buffer`s, or validate that `Buffer` input is well-formed UTF-8 before calling `replace`, `split`, or the `RE2` constructor — for example `Buffer.compare(Buffer.from(buf.toString('utf8')), buf) === 0`. Reported by [@OvOhao](https://github.com/OvOhao) in [#272](https://github.com/uhop/node-re2/issues/272).
| Version | Type | Source | Base | Exp | Impact | Vector |
|---|---|---|---|---|---|---|
| 3.1 | Secondary | GHSA | 5.1 | — | — | CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:L |
| 3.1 | Secondary | NVD | 5.1 | 2.5 | 2.5 | CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:L |
| 3.1 | Secondary | ENISA EUVD | 5.1 | — | — | CVSS:3.1/AV:L/AC:L/PR:N/UI:N/S:U/C:L/I:N/A:L |