Issue #19315 has been updated by Eregon (Benoit Daloze). Right, if one wants a NUL-terminated C string, i.e. a `char*` with no `\0` in the middle and only at the end, then `StringValueCStr()` is already the correct answer. I' believe migrating usages of `RSTRING_PTR()` which want `\0`-terminated to `StringValueCStr()` is the right thing to do. And in fact current usages of `RSTRING_PTR()` that expect `\0`-termination are incorrect, if `\0` can be in the middle of the String (though it results in a smaller C string rather than reading out of bounds). How about this: 1. Deprecate `RSTRING_PTR()` with a message saying to use either `StringValueCStr()` (if want to use it as a NUL-terminated C string) or `RSTRING_START()` (for efficiency, must be paired with either `RSTRING_LEN()` or `RSTRING_END()`). 2. Don't change the behavior of `RSTRING_PTR()` yet (so no `SHARABLE_MIDDLE_SUBSTRING`), let some time for people to migrate off of `RSTRING_PTR()`. 3. Then some release(s) later enable `SHARABLE_MIDDLE_SUBSTRING`. `RSTRING_PTR()` at that point is the same as `RSTRING_START()` (so no issues with using a different buffer than `RSTRING_END()` and already-correct usages of `RSTRING_PTR` are preserved even if they didn't migrate), but no longer guarantees to be NUL-terminated (which was already incorrect for the `\0` in the middle case and should have used `StringValueCStr()` already). I was thinking maybe we could also change `RSTRING_PTR()` to be `StringValueCStr()` in that last step, but then we get the different buffer issues mentioned above, and I think that's worse. Notably because we'd break correct usages of `RSTRING_PTR()` to help incorrect usages of `RSTRING_PTR()` (which should be `StringValueCStr()` anyway). I think we do need to deprecate `RSTRING_PTR()`, otherwise we have no way to tell existing users of `RSTRING_PTR()` that they should change to `StringValueCStr()` if they want NUL-termination (well, only `NEWS` file & docs or so, but very few will see that). ---------------------------------------- Feature #19315: Lazy substrings in CRuby https://bugs.ruby-lang.org/issues/19315#change-117474 * Author: Eregon (Benoit Daloze) * Status: Open ---------------------------------------- CRuby should implement lazy substrings, i.e., "abcdef"[1..3] must not copy bytes. Currently CRuby only reuse the char* if the substring is until the end of the buffer. But it should also work wherever the substring starts and ends. Yes, it means RSTRING_PTR() might need to allocate to \0-terminate, so be it, it's worth it. There is already code for this (`SHARABLE_MIDDLE_SUBSTRING`), but it's disabled by default and `RSTRING_PTR()` needs to be changed to deal with this. It seems a good idea to introduce a variant of `RSTRING_PTR` which doesn't guarantee \0-termination, so such callers can then use the existing bytes always without copy. There are countless workarounds for this missing optimization, all not worth it with lazy substring and all less readable: * https://bugs.ruby-lang.org/issues/19314 * https://bugs.ruby-lang.org/issues/18598#note-3 * https://github.com/ruby/net-protocol/pull/14 * Manual lazy substrings which track string + index + length * More but I don't remember all now, feel free to comment or link more urls/tickets. -- https://bugs.ruby-lang.org/