Where an array's elements are written when the visitor takes the bulk
hand-off (Visitor.arrayBulk): the destination it already owns, handed
over once, instead of one callback per element.
Why this replaced the per-element callbacks. There used to be an
arrayUnsigned / arraySigned / arrayFp32 / arrayFp64 beside this, one call
per element. Measured with bench/run_callgrind.sh's method over 1000-element
arrays (Ir/op for a message that is one array), against this hand-off filling
the same destination:
Floats gain most because their reading was already bulk (the §6.6.2 handle in
the element drain), so the callback was very nearly all that was left.
A short array is not the exception. The fixed cost — the offer, the
target's resolution, the bound's validation — is about 600 Ir, but the hand-off
takes a whole array in one call, tail elements included, so it is paid once:
four elements at the end of a 37-byte message cost 578 Ir against the 563 the
removed callbacks cost, and everything longer is the table above.
What did move is the consumer's side, for a consumer that folds. An element
callback let a reader sum, hash or convert inside the delivery; reading a
filled destination is a second pass. A reader that wants the values where they
are — which is what generated code wants — has no second pass and gets the
table. A reader that folds pays one: the same four-element array costs 1 063 Ir
if it sums the destination at arrayEnd, against 563 folding per element.
It is a cheaper pass than the calls it replaced on any array long enough to
matter, and on a four-element one it is not.
Declining costs nothing at all. An array no visitor takes is walked over
rather than decoded: array<fp64> 36 890 → 9 383 Ir/op, array<u16> 184 025 →
152 082.
This is how array elements are delivered — the only way. There is no
callback per element beside it: §5.3.1 allows a rule one implementation, and
two delivery routes for the same elements were two places for the element bound
to be compared, two resume paths to keep in step, and a standing invitation for
generated code to take the slower one. arrayBegin and arrayEnd still fire
for every array; returning null declines delivery, and the elements are then
walked over without being decoded into existence — the skip half of §6.7.2's
two intents, which is what a visitor that declares no arrayBulk gets for
every array.
The destination is filled ascending from index 0, one write per element,
and it must stay valid — same object, same length — until arrayEnd, which for
a chunked decode is several IStream.feed calls later. A plain-array
destination (values / longs) is cut to the elements written when the array
ends — including to zero for an array that is empty on the wire, and to the
prefix when an element is refused — so reusing one across arrays or messages can
never leave the previous array's tail behind, and after arrayEnd its length
is exactly this array's element count. (Cutting at the end rather than emptying
up front is measured: emptying first makes every element write a grow, which
cost array<u16> 181 583 → 231 815 Ir/op.) A typed destination is neither cut
nor emptied — it cannot be — and must already hold count elements.
A decode that fails inside an array for any other reason — malformed bytes, or
input that simply ends — leaves the destination holding what had been written
when it stopped. length is a statement about the array only once arrayEnd
has been raised or an element has been refused. This is the one
place the codec holds a reference to the caller's storage between calls (§6.6
otherwise holds nothing past a callback), and the reference is dropped at
arrayEnd and whenever a pooled machine is released.
A refused element leaves the destination holding the prefix: everything
before it is written, the offending element and everything after it is not, and
a plain-array destination is cut to exactly that length. Like every INVALID
verdict it is terminal (§5.2.1).
Where an array's elements are written when the visitor takes the bulk hand-off (Visitor.arrayBulk): the destination it already owns, handed over once, instead of one callback per element.
Why this replaced the per-element callbacks. There used to be an
arrayUnsigned/arraySigned/arrayFp32/arrayFp64beside this, one call per element. Measured withbench/run_callgrind.sh's method over 1000-element arrays (Ir/op for a message that is one array), against this hand-off filling the same destination:array<u16>into anumber[]array<u64>into aLong[]array<u64>into IntegerArrayTarget.lo/hiarray<fp64>into aFloat64Arrayarray<fp32>into aFloat32ArrayFloats gain most because their reading was already bulk (the §6.6.2 handle in the element drain), so the callback was very nearly all that was left.
A short array is not the exception. The fixed cost — the offer, the target's resolution, the bound's validation — is about 600 Ir, but the hand-off takes a whole array in one call, tail elements included, so it is paid once: four elements at the end of a 37-byte message cost 578 Ir against the 563 the removed callbacks cost, and everything longer is the table above.
What did move is the consumer's side, for a consumer that folds. An element callback let a reader sum, hash or convert inside the delivery; reading a filled destination is a second pass. A reader that wants the values where they are — which is what generated code wants — has no second pass and gets the table. A reader that folds pays one: the same four-element array costs 1 063 Ir if it sums the destination at
arrayEnd, against 563 folding per element. It is a cheaper pass than the calls it replaced on any array long enough to matter, and on a four-element one it is not.Declining costs nothing at all. An array no visitor takes is walked over rather than decoded:
array<fp64>36 890 → 9 383 Ir/op,array<u16>184 025 → 152 082.This is how array elements are delivered — the only way. There is no callback per element beside it: §5.3.1 allows a rule one implementation, and two delivery routes for the same elements were two places for the element bound to be compared, two resume paths to keep in step, and a standing invitation for generated code to take the slower one.
arrayBeginandarrayEndstill fire for every array; returningnulldeclines delivery, and the elements are then walked over without being decoded into existence — theskiphalf of §6.7.2's two intents, which is what a visitor that declares noarrayBulkgets for every array.The destination is filled ascending from index 0, one write per element, and it must stay valid — same object, same length — until
arrayEnd, which for a chunked decode is several IStream.feed calls later. A plain-array destination (values/longs) is cut to the elements written when the array ends — including to zero for an array that is empty on the wire, and to the prefix when an element is refused — so reusing one across arrays or messages can never leave the previous array's tail behind, and afterarrayEnditslengthis exactly this array's element count. (Cutting at the end rather than emptying up front is measured: emptying first makes every element write a grow, which costarray<u16>181 583 → 231 815 Ir/op.) A typed destination is neither cut nor emptied — it cannot be — and must already holdcountelements.A decode that fails inside an array for any other reason — malformed bytes, or input that simply ends — leaves the destination holding what had been written when it stopped.
lengthis a statement about the array only oncearrayEndhas been raised or an element has been refused. This is the one place the codec holds a reference to the caller's storage between calls (§6.6 otherwise holds nothing past a callback), and the reference is dropped atarrayEndand whenever a pooled machine is released.A refused element leaves the destination holding the prefix: everything before it is written, the offending element and everything after it is not, and a plain-array destination is cut to exactly that length. Like every
INVALIDverdict it is terminal (§5.2.1).