ADR 0110: Class-Variable Shadow Write-Through for Foreign NLR Relay

Status

Superseded by ADR 0130 (2026-10-06); was Implemented (2026-08-03)

Superseded by ADR 0130 (2026-10-06). Class variables now live in one place, the class process's dictionary, for the duration of a class-method invocation, so there is no shadow to write through and no relay to read it on. The shadow key, the nlr_relay shadow read, the per-scope commit map and the class_var_result return protocol were deleted. A ^ through a foreign process keeps the writes made before it because those writes were put into the one home, not because a shadow carried them. Everything below, including the amendments and the "known limit" paragraphs, is kept as history and no longer describes the code.

Implementation Tracking

Epic: BT-3035 Issues: BT-3036 (runtime outcome variant + shadow read/erase) → BT-3037 (codegen emission + BUnit regression suite) → BT-3038 (docs + e2e close-out) Status: Done

Context

Problem statement

A class method that mutates a class variable and then invokes a caller-supplied block loses that mutation if the block escapes with ^ (a non-local return, NLR):

Value subclass: CollectionDriver
  classState: runs = 0

  class countedRun: aBlock :: Block over: aList :: List -> Nil =>
    self.runs := self.runs + 1
    aList do: [:x | aBlock value: x]
    nil

  class runCount -> Integer => self.runs
Value subclass: CollectionProbe
  class escapeAfterCountedRun: aList :: List -> Integer =>
    CollectionDriver countedRun: [:x | x > 1 ifTrue: [^x]] over: aList
    0
before := CollectionDriver runCount
(CollectionProbe escapeAfterCountedRun: #(1, 2, 3))   // => 2   (correct, BT-3022)
CollectionDriver runCount                              // => before   (WRONG, want before + 1)

The returned value is correct; only the class-variable mutation is lost. This is not a hypothetical — it is reproducible today and documented as a known trap in docs/beamtalk-language-features.md § Passing Blocks Through Class Methods.

Why this happens

ADR 0041 (Universal State-Threading Block Protocol) established that every non-local return throws a 4-tuple {'$bt_nlr', Token, Value, State}, and a method's own try/catch recovers State only when the token matches its own (CatchTok =:= OurToken). For a class method, that State slot is the method's threaded ClassVars; the catch arm for a matching token yields {'class_var_result', Value, State} (NlrBoundary::ClassMethod in crates/beamtalk-core/src/codegen/core_erlang/mod.rs).

The failure case is a foreign NLR: a block invoked indirectly — through aList do: [:x | aBlock value: x], which lowers to lists:foldl/similar inside compiled library code — whose ^ belongs to a different method's frame (here, CollectionProbe escapeAfterCountedRun:). Its token does not match CollectionDriver countedRun:over:'s own token, so the catch guard fails and the exception falls through to the generic arm, which re-raises the tuple unchanged. That tuple's State field holds the foreign frame's own state (whatever CollectionProbe's method was threading), not CollectionDriver's ClassVars — the two are different buckets, and there is no slot to carry both simultaneously.

The re-thrown exception propagates to the Erlang runtime layer. apply_class_method_fun/6 (runtime/apps/beamtalk_runtime/src/beamtalk_class_dispatch.erl) already distinguishes this case from a genuine failure — a dedicated clause matches throw:Nlr:NlrST when ?IS_NLR(Nlr) and passes it through unlogged (added for BT-3022), separately from the catch-all ErrClass:Error:ErrST clause that logs and classifies as a real error. But both outcomes are then folded into the same {error, {raised, _, _, _}} shape by apply_class_method_in_context/6, and invoke_class_method/7's handler for that shape always replies with the pre-call ClassVars:

{error, {raised, _ErrClass, Error, _ST}} ->
    {reply, {error, Error}, ClassVars}     %% original, pre-call ClassVars

There is no NewClassVars value available at this point to use instead — the mutated value exists only as a let-bound Core Erlang variable inside the try body, invisible to the catch handler.

Why there is no surgical per-call-site fix

Wrapping each block-invocation site in the class method's own generated code (to catch a foreign NLR and re-attach ClassVars before re-throwing) does not generalize: in the repro above, the block is invoked indirectly, through beamtalk_collection:to_list/1 and a lists:foldl inside do:'s own implementation. Codegen has no call site there to instrument — the block escapes into library code the class method's generated body never touches directly.

Existing precedent for a write-through shadow

beamtalk_actor.erl already solves a shaped-alike problem for actor state. self_dispatch/2's unwrap_dispatch_result/1 writes the actor's current state into the process dictionary (put('$bt_actor_state', NewState)) after every nested self-dispatch call, so other code running later in the same handler sees a live, current state without it being threaded functionally through every intermediate call frame. restore_dispatch_pdict/1 runs once, in the after block of the outermost handle_call/handle_cast — it restores whatever value (or absence) preceded that call, so the shadow never leaks into the next, unrelated invocation of the same gen_server. This ADR follows the same two-part discipline — write-through during the call, single cleanup at the outer boundary — not a per-nesting-level save/restore; see the Runtime change below for why that distinction matters here.

Constraints

Decision

Add a process-dictionary shadow write-through, scoped narrowly to the one case that needs it: relaying a foreign NLR out of a class method that mutated class vars beforehand. The existing functional ClassVars threading is untouched for every other path, and the relay mechanism (class_send_dispatch/3, metaclass_send_dispatch/4, class_self_dispatch/4) is unchanged. The runtime diff is confined to one module: a new {nlr_relay, ...} variant in class_method_outcome(), produced by the two catch clauses that already distinguish a relayed ^ from a genuine error, and consumed by invoke_class_method/7 — the one place that decides what ClassVars value the class gen_server retains.

Codegen change

Class-var assignment (self.foo := value) already funnels through one function, CoreErlangGenerator::generate_field_assignment (crates/beamtalk-core/src/codegen/core_erlang/expressions.rs), which emits let ClassVarsN = call 'maps':'put'(field, Val, ClassVars{N-1}) in and flips a sticky per-method class_var_mutated flag via next_class_var() (core_erlang/mod.rs) — the same flag that already gates NlrBoundary::ClassMethod { has_class_vars }. Add one more line at that emission site, writing the just-updated map into a fixed process-dictionary key:

let Val = <value> in
let ClassVars1 = call 'maps':'put'('runs', Val, ClassVars0) in
let _ = call 'erlang':'put'('$bt_class_vars_shadow', ClassVars1) in

This must be gated on self.block_depth == 0, not just self.in_class_method(). in_class_method() is a lexical flag that stays true for the entire method body, including inside nested block literals — but a block literal written inside class P's method can be invoked from a different class's process (per ADR 0109 / the block-runs-where-invoked semantics documented in beamtalk-language-features.md). Writing the shadow unconditionally would let a block invoked while executing inside class C's gen_server write class P's ClassVars into the pdict key that C's own invoke_class_method/7 reads back — corrupting an unrelated class's persisted state on a foreign-NLR relay. block_depth (already tracked in CoreErlangGenerator, incremented in generate_block) distinguishes "top-level statement in the method body" from "inside a nested block," so the gate becomes self.in_class_method() && self.block_depth == 0 && self.class_var_names().contains(field_name).

Superseded for open classes by the BT-3667 amendment below. This scoping is not a new limitation — it matches existing behavior. generate_block already saves and restores class_var_version around a block's body (BT-1550, "so self-calls inside a conditional branch don't leak ClassVars{N} bindings into the outer scope"), meaning a class-var mutation made inside a block is already discarded on the method's normal-return path today. Scoping the shadow to block_depth == 0 keeps the shadow's contents consistent with what already survives normal return — it does not shadow anything that wasn't already going to be lost.

This only fires for top-level mutations in class methods that already have class_var_mutated = true — i.e., exactly the methods that already pay the ClassVars threading cost today. Methods with no class-var mutations, and blocks nested inside any method, emit nothing new; no new analysis pass is needed.

Amendment (2026-08-04, BT-3039): the single well-known key '$bt_class_vars_shadow' is not enough. A block literal invoked from a different class's process (exactly the case the block_depth == 0 gate above was already written to worry about) can itself contain a mutating self-send — self bump where self is the block's captured home-class identity, not the process it happens to execute in. bump's own top-frame mutation (block_depth resets to 0 on entry to its own method body) writes the shadow too, under the same global key, physically inside the foreign process — clobbering that process's own class's shadow write before its invoke_class_method/7 reads it back. The fix: key the write by the dynamic runtime identity of self for this invocation, not a single shared key:

let _ = call 'erlang':'put'({'$bt_class_vars_shadow', call 'erlang':'element'(2, ClassSelf)}, ClassVars1) in

element(2, ClassSelf) (the #beamtalk_object.class field) is used rather than a class name baked in statically at compile time: a method's self.class_name() at codegen time is fixed to the class where the method is defined, but ClassSelf — already threaded dynamically as this call's class = beamtalk_class_registry:class_object_tag(ClassName), per the Runtime change below — correctly reflects the calling class for an inherited self-dispatch chain (self otherClassMethod: inherited from an ancestor still tags its shadow write with the subclass's own identity, not the ancestor's) while also correctly separating a block's foreign captured self (tagged with its own home class) from the process it executes in. A static class-name tag would get the inherited case wrong; the dynamic ClassSelf field gets both cases right for the same reason it already exists — see the class_mod/class split noted in the Runtime change below.

ClassBuilder class methods are covered by the same emission. Runtime-installed class-method funs (ADR 0084) are not a separate implementation: generate_class_method_fun_from_block (gen_server/methods.rs) lowers each classMethods: block through the shared class-method body path, with enter_builder_class_method_context setting in_class_method = true and the cascade's classVars: keys as class_var_names — so generate_field_assignment's class-var branch, and with it the shadow write, fires inside builder funs exactly as in compiled methods. One adjustment is required for the gate to hold there: generate_class_method_fun_from_block already resets class_var_version and class_var_mutated on entry (the fun body is a fresh method frame); it must also save/reset/restore block_depth the same way, because the builder cascade is itself an expression that may lexically sit inside a block (block_depth > 0 at the cascade's position) even though the fun body executes at runtime as a class method's own top frame. Without that reset, a builder cascade written inside a block would silently lose the shadow write; with it, "block_depth == 0" uniformly means "the method's own top frame" across both compilation paths.

Scope boundary: the one path not covered is a class method implemented directly in hand-written Erlang that mutates class vars by returning {class_var_result, ...} itself. As of this ADR, no such method exists — no module under beamtalk_stdlib/src or beamtalk_runtime/src produces class_var_result (only the dispatch layer and beamtalk_supervisor consume it) — so this is an FFI authoring rule, not a live gap: an Erlang-implemented class method that mutates class vars and can have a foreign NLR pass through it should also put('$bt_class_vars_shadow', NewCV) at its mutation points. Record this in docs/development/erlang-guidelines.md.

Runtime change

The relay-vs-genuine-error distinction is already made in exactly two places — the throw:Nlr:NlrST when ?IS_NLR(Nlr) catch clauses of apply_class_method_fun/6 and apply_compiled_class_method/7 (both added for BT-3022) — and then immediately erased by folding both into the same {error, {raised, ...}} shape. Instead of re-inferring the distinction downstream from tuple shape (fragile: it would silently break if those catch clauses were ever consolidated), carry it in the outcome type. Those two clauses change from {error, {raised, throw, Nlr, NlrST}} to a new variant:

-type class_method_outcome() ::
    test_spawn
    | {ok, term()}
    %% BT-3032: a foreign `^` passing through — control flow to relay, not a failure.
    | {nlr_relay, term(), list()}
    | {error, #beamtalk_error{}}
    | {error, undef_in_body}
    | {error, {raised, atom(), term(), list()}}.

invoke_class_method/7 handles the new variant by reading the shadow, and wraps everything in try ... after so the shadow is erased on every path:

invoke_class_method(Selector, Args, ClassName, _Module, DefiningClass, DefiningModule, ClassVars) ->
    try
        case apply_class_method_in_context(Selector, Args, ClassName, DefiningClass, DefiningModule, ClassVars) of
            test_spawn ->
                test_spawn;
            {ok, {class_var_result, Result, NewClassVars}} ->
                {reply, {ok, Result}, NewClassVars};
            {ok, Result} ->
                {reply, {ok, Result}, ClassVars};
            %% BT-3032: a foreign `^` relay is not a failure of *this* method —
            %% recover ClassVars mutated before the relay from the shadow written
            %% by generate_field_assignment, instead of reverting to ClassVars
            %% as it stood before this call. The reply shape is unchanged, so
            %% class_send_dispatch/3's existing `{error, Nlr} when ?IS_NLR(Nlr)`
            %% re-throw clause (BT-3022) works untouched.
            {nlr_relay, Nlr, _ST} ->
                %% BT-3039 amendment: keyed by this call's own class tag so a
                %% foreign class's shadow write (see amendment above) can never
                %% be read back here.
                ShadowKey = {'$bt_class_vars_shadow', beamtalk_class_registry:class_object_tag(ClassName)},
                NewClassVars =
                    case erlang:get(ShadowKey) of
                        undefined -> ClassVars;   %% no mutation occurred before the relay
                        Shadow -> Shadow
                    end,
                {reply, {error, Nlr}, NewClassVars};
            {error, #beamtalk_error{} = Error} ->
                {reply, {error, Error}, ClassVars};
            {error, undef_in_body} ->
                {reply, {error, undef}, ClassVars};
            {error, {raised, _ErrClass, Error, _ST}} ->
                %% Genuine error: revert, exactly as today.
                {reply, {error, Error}, ClassVars}
        end
    after
        erlang:erase({'$bt_class_vars_shadow', beamtalk_class_registry:class_object_tag(ClassName)})
    end.

The other consumer of class_method_outcome(), unwrap_self_dispatch_outcome/3, gains a clause with behavior identical to how those throws unwind today — currently there is no throw-specific clause; the generic {error, {raised, ErrClass, Error, ST}} -> erlang:raise(ErrClass, Error, ST) catch-all handles them, binding ErrClass = throw for a relayed NLR:

{nlr_relay, Nlr, ST} ->
    erlang:raise(throw, Nlr, ST);

The explicit variant means dialyzer sees the relay case as part of the contract: a future refactor of either apply function's exception handling cannot silently collapse relay into error — the variant would have to be deliberately removed, which is loud, not silent.

The after clause — not a bare trailing statement — is required for the erase: apply_class_method_in_context/6 has an un-try'd prelude (beamtalk_class_registry:class_object_tag/1, lookup_class_method_fun/2, is_test_execution_selector/1) that could in principle raise before the case ever completes, and the class gen_server is long-lived, so any path that skipped the erase would leave a stale shadow for the next unrelated call to read. try ... after ... end guarantees the erase runs regardless.

This guarantees the shadow never survives past a single class_method_call/metaclass_method_call, so a later, unrelated call on the same class gen_server can never read a stale value. invoke_class_method/7 is always the outermost frame for a given external call: class_self_dispatch/4 and class_self_dispatch_local/4 (the self-send path) call apply_class_method_in_context/6 directly and run in the same process without a new gen_server hop, so a chain of self-dispatched calls within one external call writes the shadow sequentially (each mutation overwrites the previous value) and a foreign NLR relayed through any depth of that chain still unwinds to this same try before the call returns. No per-nesting-level save/restore is needed — unlike beamtalk_actor.erl's '$bt_actor_state', which restores the pre-call value in its after because instance self-dispatch exposes "the currently live state" during the chain for other code to read; here nothing reads the shadow until after the whole chain has finished, so straight overwrite-and-erase-once-at-the-end is sufficient. Cross-class self-dispatch is impossible by construction (class_self_dispatch only walks one class's own superclass chain), and different classes are different gen_server processes with independent process dictionaries, so there is no cross-class shadow collision through self-dispatch when the process a class method runs in is that class's own gen_server.

Amendment (2026-08-04, BT-3039): that premise doesn't hold for a block. A block literal is lexically part of one class's method but, per ADR 0109, executes in whichever process invokes it — which can be a different class's gen_server entirely. A mutating self-send inside such a block (self bump, self's captured home class) then runs class_self_dispatch's target method body physically inside the foreign process, and that body's own top-frame shadow write (see the Codegen change amendment above) was landing in the single global key — the exact collision this paragraph said couldn't happen, just reached by a different route than a literal cross-class class_self_dispatch call. The class-keyed shadow (Codegen change amendment) closes this: the foreign write is now tagged with its own class's identity, so it can never be read back by the process's own invoke_class_method/7, which only ever reads its own class-tagged key. The foreign entry itself is simply never read by the process's own invoke_class_method/7 — a bounded, harmless stale process-dictionary entry (bounded by the number of classes that have ever passed a mutating block into this one, i.e. by the total class count in the running system) that lives until this gen_server restarts. See Consequences below.

Amendment (BT-3667, revised by BT-3675): BT-3667 let a class-side override's class-variable write survive scopes that cannot thread a ClassVars rebind out lexically (a bare block, a loop body, a conditional or match: arm, an on:do:/ensure: body, a send whose prelude is closed into an opaque value) by re-reading this shadow from generated code. That was unsound: the shadow is written when a callee writes, so a write by a callee that then raised, and was caught, was brought back by the next compiled read, violating this ADR's own "a genuine runtime error after a class-var mutation must still revert the mutation". BT-3675 removes every compiled-code reader of the shadow. Compiled code no longer reads it at all; the premise above (nothing but invoke_class_method/7's NLR relay reads it, and only after the whole chain finished) holds again.

The mid-chain recovery uses per-scope commits instead, keyed by a per-scope token, in a second process-dictionary entry ('$bt_class_vars_commit', a #{Token => ClassVars} map) that is only touched through beamtalk_class_dispatch:class_var_scope_commit/3, class_var_scope_read/3, class_var_scope_take/3 and class_var_scope_export/3:

Known limit: a closure stored in a local and invoked by a later statement does not keep its class-variable writes. b := [self foo]. b value. self noop (and b value. b value) drop the write, whether foo is a subclass override or the class's own mutating selector (b := [self bump]; both are pinned). The compiler warns about this shape (BT-3681, @expect stored_closure; see the amendment below). The write is dropped because the closure exports into the scope of the statement that built it, which that statement's refresh already consumed, and a later statement has no scope of its own. BT-3667's shadow read happened to keep it, which is exactly the resurrection BT-3675 removes. An attempt to remember such scopes for the rest of the body was withdrawn: it leaked unbound token names out of closures without an enclosing scope, and it could not keep a twice-invoked closure correct without recency stamps on every commit and on the lexical copy. The closure's own statements still behave normally, a writing send in the same body is kept, and the behaviour is pinned in self_send_override_blocks_test.bt (testStoredClosure…). A closure invoked in the same statement that builds it by a stdlib higher-order method or construct ([self foo] value, collect:, do:, ensure:, on:do:, Result tryDo:) keeps its writes.

A stored closure can also read a stale value, or one resurrected after a caught raise: the building statement's token stays in the closure's sync chain after that statement's refresh consumed it, so a later invocation can start from an older entry than the lexical copy. b := [self nextId]. [b value. 1 / 0] on: Error do: […]. id := b value answers 2 instead of 1 (the raised try body's write, which the catch should have discarded, is read back: it is resurrected, the same bug class BT-3675 fixes for direct sends), and b value. self bump. b value answers 2 instead of 3 (the second call reads the older, stale entry). Both are pinned (testStoredClosureReads…) and covered by the BT-3681 diagnostic.

Known limit: a block literal passed to a user-defined class-side higher-order method does not keep its class-variable write. With class section: aBlock => aBlock value, self section: [self foo] drops a subclass override's write: the mechanism differs by position. At a method's top level the statement is a precise class-side send with no scope of its own, so the block's rebind is confined to its fun and section: returns the ClassVars it was given. Inside a confined scope (a loop body, an arm, an on:do: body) the closure does export its write into the scope's token, but the send's own commit after it returns (the stale ClassVars it was called with) overwrites that export (pinned for a non-writing HOM in a loop body; a HOM that writes a class variable itself is a compile error inside a bare block). The HOM's own writes are kept (class timed: aBlock => self.calls := self.calls + 1. aBlock value called as self timed: [self noop] leaves calls incremented), which is why this was left as a limit: making the send its own scope loses the HOM's own write instead, unless the passed, returned and exported class variables are merged three ways at the refresh, which is a larger change. Both limits are pinned in self_send_override_blocks_test.bt; the diagnostic is BT-3681 (see the amendment below) and the merge is tracked in BT-3682, which must cover both positions.

Amendment (BT-3681): the stored-closure limits are diagnosed at compile time. The compiler (semantic_analysis/validators/stored_closure_validators.rs) now emits a warning, category stored-closure (suppress with @expect stored_closure, or stored-closure = "off" in [diagnostics]), for two shapes in a class method: a block bound to a local (b := [self foo]) whose body makes a class-side self/own-class send that may write a class variable, when a later statement of the same body or block invokes the local (b value, value:, value:value:, …), which covers the stale and resurrected reads too, since the same send is what reads them; and a block literal passed to a user-defined class-side higher-order method (self section: [self foo]). A send "may write" when the selector resolves to a user-defined class method and: for a method the class defines, compute_class_var_mutating_selectors (the rule the codegen gates use, reused, not copied) cannot prove it pure, or, for a self send, the class and the method are not sealed (a subclass override may write, the ADR's own case; an explicit Counter foo binds directly, so only the first applies); for a user-defined method the class only inherits, always (that rule's "not defined locally, assume the worst" call, whatever the sealing). This is deliberately conservative: a pure sealed method inherited from a user-defined parent that is not in the same module is still flagged (the class cannot see the parent's body; BT-3688 narrowed this, see below), and @expect stored_closure suppresses such a case. Stdlib class methods and unresolvable selectors are never flagged, and a class with no class variable (own or inherited) is skipped. Never flagged: [self foo] value, collect:/do:/ensure:/on:do: and other blocks invoked by the statement that builds them, a stored block that is never invoked by a later statement, and a stored block rebound before it is invoked (b := [self foo]. b := [0]. b value). A stored block handed by name to a user-defined class-side higher-order method (self section: b) counts as invoked; the two gaps that sentence originally listed (a stdlib higher-order method, a conditional rebind) are dispositioned by the BT-3688 amendment below. The warning carries the hint to invoke the block in the statement that builds it or to return the value and write the class variable from the method body, and a note at the invocation. It is a warning, not the error the analogous self.field := in a stored closure gets, because the write happens through a send and whether it writes is only known at run time. The runtime behaviour is unchanged and stays pinned in self_send_override_blocks_test.bt, whose fixtures carry @expect stored_closure; the alternative (recency stamps on commits so the write is kept) was not pursued.

Amendment (BT-3688): stored-closure advisory false positives and gaps. Fixed, each with a test in semantic_analysis/tests/bt3681_stored_closure_class_var_writes.rs: (1) a later block that declares the stored local's name as a parameter (b := [self bump]. #(1) do: [:b | b value]) no longer counts as a use (the search skips that block only; a use outside it still does), and neither do the statements of a nested block that follow a rebind of the local in that block's own sequence; (2) a stored block handed by name to a collection higher-order method (items do: b, collect:, select:, detect:, inject:into:, ... the table opaque_fold_callable_arg that the ADR 0128 fold already shares, reused, not copied; a self/super receiver is excluded the way is_opaque_callable_hom_send excludes it, since that dispatches to a possibly user-defined method) is warned like a direct invocation (self_send_override_blocks_test.bt pins at runtime that the write is lost there too); (3) standalone Foo class >> bar => ... definitions are scanned as class methods of Foo (package-qualified pkg@Foo class >> ones are not: the hierarchy lookup is by simple name); (4) a pure class sealed method inherited from a parent in the same module is no longer flagged: "may write" is now judged by compute_class_var_mutating_selectors run on the defining class, whose body is visible, unless that inherited method makes a self send (it late-binds to the receiving class, whose override may write: Base class sealed pure => self helper reached from a Leaf that overrides helper to write stays flagged, open or sealed Leaf). Closed as intended, pinned by tests: an invocation after a conditional rebind (cond ifTrue: [b := [0]]. b value) still warns, since when the branch is not taken b is still the stored block and the check does not prove every branch rebinds; and a method whose defining class's body is not visible (a parent in another file or package, or a method added by a standalone definition) is still assumed to write (@expect stored_closure suppresses). A stored block handed to a collection method outside that shared table, or to on:do:/ensure: by name, is still not tracked.

Amendment (BT-3683): a scope's refresh falls back to the enclosing scopes' newest commit, not the lexical version before it. A scope's refresh takes only its own token. When nothing committed under it (in practice an arm that was not taken), the fallback used to be the lexical ClassVars version live before the scope. Inside a loop body or closure that version is the stale copy captured at the start of the enclosing scope, so #(1, 2) do: [:i | i =:= 1 ifTrue: [self increment]. i =:= 2 ifTrue: [self nonWriting]] lost the write of iteration 1: iteration 2's skipped first arm refreshed to the stale copy and committed it over the enclosing loop scope's newer entry. The fallback is now class_var_scope_read(<enclosing tokens>, <lexical version>): the newest value any enclosing scope committed, and the lexical version only when none did. Every ClassVars version minted inside a scope (a send, a nested refresh, a direct write, a threaded construct's rebind) is committed to the innermost token, so an enclosing commit is never older than the lexical version within one iteration. A skip-no-op-commits optimisation (not committing when the callee returned the ClassVars it was given) was tried and lost writes for the same reason; the refresh depends on every mint being committed, so it was not taken. A scope that raised still leaves nothing readable: its writes are committed under its own (dead) token or the raised body's token, never under an enclosing one. Pinned in self_send_override_blocks_test.bt for open and sealed classes (do: over a literal, to:do:, timesRepeat:, ifTrue:ifFalse:, nested conditionals, on:do: bodies and handlers, a direct write between iterations, a raised iteration).

Amendment (BT-3690): a send commits only a reply that may have changed the class variables. The send's commit (class_var_scope_commit/3, after the rebind) is now made inside the {class_var_result, _, _} arm of the reply: let _ = case _CMR of <{'class_var_result', _CW, _}> -> commit(...) <_> -> 'ok' end. A plain reply means the callee returned without touching the class variables, so the rebind is the value the pre-call sync read from the scope chain (the newest commit of the enclosing scopes, else the lexical version); committing it again only re-stored what the chain already answers, for a process-dictionary read-modify-write per send (about 100 ns of the ~250 ns a confined open-class send paid). The pre-call read, the token, the closure export and every other mint's commit are unchanged, and a skipped commit leaves the chain exactly as it was: the BT-3683 invariant ("every ClassVars version minted inside a scope is committed to the innermost token") still holds, because a plain reply mints no new value. The earlier skip-no-op-commits attempt failed on the old refresh fallback (the lexical version), not on skipping itself; it is only sound with BT-3683's fallback, and the tests that catch the old combination (armsInListDo, armsDirectWriteBetween, self_send_plain_reply_test.bt) fail against it. One shape keeps committing a plain reply: a send with a block literal among its arguments (self section: [self foo]). The callee may invoke the block, whose export lands in the scope while the callee runs, and the send's own commit has to overwrite it exactly as before (the pinned BT-3682 limit, testBlockPassedToClassSideHomInLoopBody); skipping there flips that pinned answer. What the skip guarantees is only this: a plain reply from a send without a block literal argument leaves the scope chain as the pre-call sync read it. It does not guarantee that no write reaches the chain while the callee runs. A closure that reaches the callee by name (a block parameter: [:b | self section: b] value: [self foo] in a loop body) is not a block literal of the send; its export into the scope lands while section: runs, and the send, replying plain, no longer overwrites it with the stale ClassVars it was called with. That write is now kept, where self section: [self foo] still drops it (the BT-3682 limit): the two spellings of one call answer differently until BT-3682 merges the passed, returned and exported class variables. The new answer is the correct one, and it is pinned (SelfSendPlainReplyTest>>test{Sub,Sealed}ViaSectionParamInLoop answer 1, ...ViaSectionLiteralInLoop answer 0). A closure stored in a local and invoked by a later statement already exports into a consumed scope (the stored-closure limit above), so that case is unchanged. Sealed and open classes share the rule. self_send_plain_reply_test.bt pins plain sends interleaved with writing sends in loops, nested arms, raised scopes, handler arms, ensure:, closures and a late-bound foo that a subclass override turns into a write, for the open class called directly, through a subclass, and the sealed class.

Known gap found while writing those tests (BT-3691, present before BT-3690): in a sealed class, a threaded fold loop with one arm that writes a class variable and another arm that sends a provably pure selector loses the write, because the pure send mints no ClassVars version and the fold's stale accumulator is committed over the arm's export. Pinned as testSealedPinBug....

Cost: only sends inside such a scope pay (a token bind, a pre-call read, a commit (BT-3690: only for a class_var_result reply) and the scope's refresh, about four process-dictionary operations, three for a plain reply); the sealed-class and non-confined hot path pays nothing, and the BT-3667 pre-call shadow read that every late-bound open-class self-send paid is gone. A known residual: a raised scope's token entry stays in the map until the outermost dispatch returns, so a loop that raises and catches a million times in one external call holds a million small entries until it ends.

REPL example

st> CollectionDriver runCount
0
st> CollectionProbe escapeAfterCountedRun: #(1, 2, 3)
2
st> CollectionDriver runCount
1

Error example (genuine error still reverts)

class countedRun: aBlock :: Block over: aList :: List -> Nil =>
  self.runs := self.runs + 1
  1/0.   "genuine error after mutation"
  nil
st> CollectionDriver runCount
0
st> CollectionDriver countedRun: [:x | x] over: #(1)
Error: ArithmeticError: division by zero
st> CollectionDriver runCount
0    "mutation reverted, exactly as today"

Prior Art

Pharo/Squeak Smalltalk — why this bug class doesn't exist there

This bug is BEAM-specific, not Smalltalk-general, and it's worth being explicit about why. In Pharo, a class variable is a true mutable slot shared by the class and its instances; a non-local return unwinds the real call stack (via BlockContext/thisContext machinery) directly to the home context, and there is no separate "commit" step for a class-variable write to survive — the assignment already happened, in place, before the unwind. Not adopted, but explains the constraint: Beamtalk inherits Smalltalk's mutable-class-variable semantics while compiling to an immutable substrate (Core Erlang has no mutable variables), so it must reconstruct "the assignment already happened" via functional threading — and it's precisely the threading reconstruction that this bug exposes as incomplete for one relay path. The fix's goal is to restore the Pharo-equivalent guarantee, not to add a new one.

Kotlin/Java — mutable fields, same story as Smalltalk

A companion object field in Kotlin (or a static field in Java) mutated before a lambda-escaping early return (e.g. via a labeled return@outer or a checked exception used for control flow) is never "lost," for the same reason as Pharo: the write is a direct memory mutation, not a functional rebinding that a catch handler might fail to recover. Confirms the same point from the mainstream-OO side: this entire bug class is an artifact of choosing functional state-threading as the compilation strategy for mutable semantics, not something inherent to "a variable mutated before a non-local return."

Erlang/OTP — process dictionary as an escape hatch, not a primary mechanism

Erlang style guides discourage the process dictionary for general state, but OTP itself uses it for exactly this shape of problem: logger metadata, seq_trace, and stdlib's own error_logger all use put/get to make state visible across call boundaries a functional threading model can't reach without threading it through every intervening function signature. Adopted: using the process dictionary as a narrow, single-purpose escape hatch rather than a general state mechanism — the functional threading remains the source of truth for every path except the one it structurally cannot cover.

Haskell — IORef as an escape from pure threading

Haskell's State monad (cited in ADR 0041 as the model for StateAcc) is the general case; IORef/STRef exist precisely for the cases where purely functional threading can't reach across a boundary (callback into foreign code, FFI). Adopted: the same split — functional threading for the general case, a mutable cell for the one boundary case that needs it.

This codebase — beamtalk_actor.erl's '$bt_actor_state'

Already discussed above under Context; this ADR extends the same pattern to class variables and follows its cleanup discipline (restore_dispatch_pdict/1's use of erase/1 in an after-equivalent path).

User Impact

Newcomer

No visible syntax change. The trap described in beamtalk-language-features.md § Passing Blocks Through Class Methods disappears rather than needing to be learned. A newcomer writing a class method that mutates state and takes a block simply gets correct behavior.

Smalltalk developer

Restores the expected Smalltalk invariant that a non-local return is ordinary control flow, not a special case that silently drops side effects — matching how ^ behaves everywhere else in the language (ADR 0041 already made this true for actor and value-type state; this closes the one remaining gap for class-side state).

Erlang/BEAM developer

The process-dictionary shadow is a narrow, well-precedented pattern (see beamtalk_actor.erl) rather than a novel mechanism — reviewers familiar with OTP idioms for crossing callback boundaries will recognize it immediately. The erase/1 discipline means it introduces no new leak surface.

Production operator

No change to the hot dispatch path for the majority of class methods (no class vars, no new codegen emitted). For the minority that do mutate class vars, cost is one extra put/2 per mutation — negligible relative to the existing maps:put/3 threading cost they already pay. No new failure mode: genuine errors keep today's revert behavior exactly.

Tooling developer (LSP, IDE)

No AST or type-level change; this is purely a runtime/codegen fix for existing, already-typed constructs. No new diagnostics are introduced (a static "this method mutates class vars and takes a Block" warning was considered in the original issue and rejected as producing more false positives than value, since that shape is common and legitimate — only ^-through-the-block is rare).

Steelman Analysis

Best argument for Option B (full write-through — class vars always live in the shadow)

CohortTheir strongest argument
BEAM veteran"One mechanism is simpler than two. If class vars always live in the process dictionary, there's no functional/shadow split to keep in sync — invoke_class_method/7 just reads the pdict, period."
Language designer"This removes an entire category of 'did the shadow get written before the read' bugs. A single source of truth is more robust than two representations of the same state that must agree."
Operator"Fewer code paths to reason about during an incident — one state model for class vars, not a functional one for the common case and a shadow for the rare one."

Why Option C (scoped shadow) wins despite the steelman

  1. Option B's runtime side is cheap, but its codegen side is not. To be honest about it: revert-on-error would be nearly free under Option B too — the pre-call ClassVars is already a bound variable at the dispatch choke point, so "restore on genuine error" is just replying with it, and the per-call put/get cost is noise next to the gen_server hop every class-method call already pays. The real cost of B is the codegen migration: every class-var read and write site, in compiled methods and ClassBuilder funs alike, changes representation, and the entire threading machinery (ClassVars{N} versioning, class_var_result unwrapping at self-send sites, the class-method NLR state slot) has to be dismantled or kept in sync during the transition. That is an L-sized refactor of working, tested code for the same observable fix.
  2. Option B has a semantics edge Option C doesn't touch. A block created in class P's method that reads a class var currently captures the threaded map value — a creation-time snapshot that stays correct wherever the block later runs. Under pdict-resident class vars, a naive read-compiles-to-get would read whatever process the block happens to execute in (wrong class, or no class at all). Preserving today's snapshot semantics for block-captured reads is solvable but is exactly the kind of subtle, cross-process regression surface this narrow bug doesn't justify opening.
  3. Smaller blast radius. Option C is one new outcome variant, two one-line catch-clause changes, one new clause plus a try/after in invoke_class_method/7, one clause in unwrap_self_dispatch_outcome/3, and two codegen touch-points (the emission line, and a block_depth reset in the builder-fun lowering) — all in existing functions, none changing how state is represented. Option B changes how every class method's state is represented at runtime (BT-3032's own investigation already flagged this cost for what it called "Option 2").

Tension point

BEAM veterans and language designers reasonably prefer Option B's conceptual simplicity; the deciding factor is that Option C gets the same observable fix with a fraction of the changed surface and none of the risk to the error-revert invariant, which the acceptance criteria treat as non-negotiable.

Alternatives Considered

Alternative A: Accept the limitation, document only

Promote the existing beamtalk-language-features.md note to a fuller worked example; close BT-3032 without a runtime fix.

Rejected: the pattern (mutate a class var, then hand a block to a method that invokes it indirectly) is a natural way to write a Collection subclass's do: delegating to a class-side helper — exactly the shape docs/beamtalk-language-features.md itself recommends elsewhere in the same section. Silent data loss in class state is a trust-eroding correctness bug even though the trigger is narrow; a fix that costs only the mutating class methods is affordable enough not to accept the limitation.

Alternative B: Full write-through (class vars always live in the shadow)

See Steelman Analysis above. Rejected not for runtime cost (revert-on-error and the per-call put/get are both cheap at the dispatch choke point) but for migration size and semantics risk: it replaces the entire class-var threading representation across compiled methods and ClassBuilder funs, and it opens a block-captured-read semantics edge (creation-time snapshot vs. executing-process pdict) that Option C never touches. It would, however, also cover hand-written Erlang class methods for free and delete the threading machinery long-term — if class-var-heavy code becomes common enough that the threading machinery itself is a maintenance burden, B is the direction to revisit.

Alternative D: Per-block-invocation-site wrapping

Wrap each place a class method's generated code invokes a block, catching a foreign NLR there and re-attaching the current ClassVars before re-throwing.

Rejected (already investigated in the BT-3032 issue itself): does not generalize. The repro's block reaches aBlock value: x through beamtalk_collection:to_list/1 and lists:foldl inside do:'s own compiled implementation — there is no call site in the class method's own generated code to instrument. Any indirect invocation through library code defeats this approach.

Alternative E: Compile-time diagnostic

Warn when a class method both mutates a class variable and takes a Block parameter.

Rejected (already investigated in the BT-3032 issue itself): that shape is common and legitimate (most block-taking class methods with class-var mutations never have the block's ^ escape); the warning would be wrong almost every time it fired.

Alternative F: Wrap the continuation at each mutation site (pure codegen, no pdict)

Instead of a process-dictionary shadow, wrap the continuation after each top-level class-var mutation in a try/catch that intercepts any escaping exception, and — if it's a non-matching (foreign) NLR — re-throws it carrying the in-scope ClassVarsN as an auxiliary payload rather than the original tuple's own (unrelated) state field:

let ClassVars1 = call 'maps':'put'('runs', Val, ClassVars0) in
try
    <rest of method body, using ClassVars1>
catch <Cls, Err, Stk> ->
    case {Cls, Err} of
      <{'throw', {'$bt_nlr', _T, _V, _S} = Nlr}> when 'true' ->
          primop 'raw_raise'('throw', {'$bt_nlr_with_cv', Nlr, ClassVars1}, Stk)
      <_> when 'true' ->
          primop 'raw_raise'(Cls, Err, Stk)
    end
end

This is a genuinely strong alternative: it never touches the process dictionary, so it has none of Option C's cross-process contamination risk (this ADR's own codegen fix above exists only because the pdict approach needed one) and composes with nesting for free (each mutation's try wraps a strictly smaller continuation than the last, so the innermost one to fire carries the most current ClassVarsN). The outermost wrap_class_method_body_with_nlr_catch would need to recognize the new {'$bt_nlr_with_cv', Nlr, CV} wrapper and use CV instead of ClassVars when relaying Nlr onward.

Rejected in favor of Option C, but noted as the strongest alternative found during review: it requires new try/catch scaffolding at every class-var mutation site (not just a one-line put/2), plus a second wrapper shape the relay path and every consumer of the NLR tuple must now recognize alongside the plain 4-tuple — more codegen surface than Option C for the same observable fix. If Option C's process-dictionary side channel proves troublesome in practice (e.g. the coupling noted in Consequences below), this is the fallback to revisit.

Consequences

Positive

Negative

Neutral

Implementation

Affected components: codegen (crates/beamtalk-core/src/codegen/core_erlang/expressions.rs — generate_field_assignment; gen_server/methods.rs — generate_class_method_fun_from_block) and runtime (runtime/apps/beamtalk_runtime/src/beamtalk_class_dispatch.erl only).

  1. Add the '$bt_class_vars_shadow' put/2 emission immediately after the existing let ClassVarsN = call 'maps':'put'(...) in in generate_field_assignment's class-var branch, gated on self.in_class_method() && self.block_depth == 0 && self.class_var_names().contains(field_name) — the block_depth == 0 clause is new; the rest is the existing condition. No new analysis pass needed.
  2. In generate_class_method_fun_from_block, save/reset/restore block_depth alongside the existing class_var_version/class_var_mutated resets, so the gate in step 1 fires correctly inside ClassBuilder class-method funs regardless of where the cascade lexically sits.
  3. Add the {nlr_relay, term(), list()} variant to class_method_outcome(); change the two throw:Nlr:NlrST when ?IS_NLR(Nlr) catch clauses (apply_class_method_fun/6, apply_compiled_class_method/7) to produce it.
  4. In invoke_class_method/7: wrap the case in try ... after erlang:erase('$bt_class_vars_shadow') end and add the {nlr_relay, Nlr, _ST} clause (shadow read, reply {error, Nlr} with the recovered class vars). In unwrap_self_dispatch_outcome/3: add {nlr_relay, Nlr, ST} -> erlang:raise(throw, Nlr, ST).
  5. Record the FFI authoring rule (Erlang-implemented class methods that mutate class vars must shadow-write) in docs/development/erlang-guidelines.md, and update the beamtalk-language-features.md § Passing Blocks Through Class Methods caveat to reflect the fix.
  6. Regression tests in stdlib/test/: the repro from this ADR's Context section; a companion test asserting a genuine error after a mutation still reverts; a third asserting a self-dispatched (self otherClassMethod:) inherited method's mutation also survives a foreign-NLR relay; a fourth asserting a mutation made inside a block passed to another class's method still behaves as it does today (discarded on normal return, not newly preserved) — locking in the block_depth == 0 scoping as intentional (superseded for open classes by the BT-3667 amendment: block-interior mutation is now kept on normal return, and that test was flipped); and a fifth running the Context repro against a ClassBuilder-defined class (including one defined inside a block) to pin the builder-fun coverage and the step-2 block_depth reset.

References