Issue #21795 has been updated by Eregon (Benoit Daloze). I thought more about `node_id` and I found at least one case where it is problematic with different versions of Prism. The problem is the `node_id` from the bytecode is computed by the builtin Prism parser used by `prism_compile.c`, while the usage of e.g. `Prism.find` might use a more recent Prism. Here is a reproduction showing the problem: ```ruby if false # A code snippet which generates a different number of nodes on Prism 1.2.0 and 1.9.0 case 1 in 2 A.print message: in 3 A.print message: end end def a end def b end require "prism" p Prism.find(method(:b)) ``` ruby master is fine: ``` $ ruby -v find_check.rb ruby 4.1.0dev (2026-03-27T16:16:27Z revert-source_loca.. f510d4103e) +PRISM [x86_64-linux] @ DefNode (location: (14,0)-(15,3)) ├── flags: newline ├── name: :b ``` But on Ruby 3.4.5 it's broken: ``` # Use latest prism, there is no prism release with Prism.find yet: $ cd prism $ bundle exec rake compile $ bundle exec ruby find_check.rb @ DefNode (location: (11,0)-(12,3)) ├── flags: newline ├── name: :a ``` This returns method `a` and not `b`! And if I change `p Prism.find(method(:b))` to `p Prism.find(method(:a))`, then ruby-master is correct, but 3.4.5 returns an `IfNode`. IOW, 3.4.5 returns the node before the correct one, which can be any node. I think this illustrates well that `node_id` is brittle, it depends on the number of nodes before the node of interest. OTOH, start line/column + end line/column (or equivalently, start & end offsets) is far more robust, because it represents an actual position in the source file, independent of the Prism version. The only confusion there would be if there are nodes with exactly the same start & end offsets, and they can't be differentiated based on the input. That's not a problem for the 5 methods proposed in this issue: This relates to the discussion [here](https://bugs.ruby-lang.org/issues/6012#note-47) with @mame about whether `node_id` is better, but from this finding it's clear `node_id` is worse if the version isn't fixed. I'll quote here the relevant part about `source_location` being able to locate the right node:
However, `source_location` is not an appropriate key to look up the AST subtree corresponding to a Ruby object.
I believe it is though, with the knowledge of what kind of node we are looking for. For example in `def foo; bar; end`, `bar` in the AST is covered exactly by both a `StatementsNode` and a `CallNode`. If we are using `Thread::Backtrace::Location#ast` we'd want the location of the call to `bar`, so we know we want the `CallNode`, not the `StatementsNode` and there is no ambiguity. I believe the same holds for all 5 methods proposed in #21795. My intuition there is all nodes listed in https://bugs.ruby-lang.org/issues/21795#note-2 cannot have the exact same `source_location` (e.g. we cannot have code with the same starting and ending position that is two of `DefNode`, `LambdaNode`, `ForNode`, `Call*Node`, `Index*Node`, `YieldNode`). Do you have a counter-example where this wouldn't hold?
No counter-example has been found. ---------------------------------------- Feature #21795: Methods for retrieving ASTs https://bugs.ruby-lang.org/issues/21795#change-116931 * Author: kddnewton (Kevin Newton) * Status: Open ---------------------------------------- I would like to propose a handful of methods for retrieving ASTs from various objects that correspond to locations in code. This includes: * Proc#ast * Method#ast * UnboundMethod#ast * Thread::Backtrace::Location#ast * TracePoint#ast (on call/return events) The purpose of this is to make tooling easier to write and maintain. Specifically, this would be able to be used in irb, power_assert, error_highlight, and various other tools both in core and not that make use of source code. There have been many previous discussions of retrieving node_id, source_location, source, etc. All of these use cases are covered by returning the AST for some entity. In this case node_id becomes an implementation detail, invisible to the user. Source location can be derived from the information on the AST itself. Similarly, source can be derived from the AST. Internally, I do not think we have to store any more information than we already do (since we have node_id for the first four of these, it becomes rather trivial). For TracePoint we can have a larger discussion about it, but I think it should not be too much work. In terms of implementation, the only caveat I would put is that if the ISEQ were compiled through the old parser/compiler, this should return `nil`, as the node ids do not match up and we do not want to further propagate the RubyVM::AST API. The reason I am opening up this ticket with 5 different methods requested in it is to get approval first for the direction, then I can open individual tickets or just PRs for each method. I believe this feature would ease the maintenance burden of many core libraries, and unify otherwise disparate efforts to achieve the same thing. -- https://bugs.ruby-lang.org/