From 17c715524850a45178b2393619378c26201f524a Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 15:21:39 -0700 Subject: [PATCH 1/4] docs(C1): a particle that is also a credential counts in the run (#554) Settles #554 as option (a), the shipped behavior; no parse moves. After a comma, a word of both the particle and the unambiguous suffix vocabulary (vd, mc) counts as a suffix word in C1's run test, while S2's company still refuses to let it speak for the word behind it. These are two questions, not one answered twice: run membership asks what the word is, company asks whether a particle can vouch for the word P2 joins it to. Written in one case, where no capitals speak, the part reads as the credential run (JOHN SMITH, VD MA -> suffix VD MA, reported), which is the reading a reader takes; (b) would give that input a one-word listing form with no given name. rules.md#C1 gains the clause and two examples; decisions.md#C1 records the decision. The examples enter the rules corpus, so the #544 run rule gains both names in the 2.0.0-2.3.0 ledgers and the moved corpus claims are re-recorded. The gate exits 0 at every baseline. Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 3 ++ docs/design/rules.md | 7 ++++ tests/v2/test_ledger_guards.py | 34 +++++++++++++++----- tools/differential/corpus_rules.jsonl | 2 ++ tools/differential/expected_since_2.0.0.toml | 7 +++- tools/differential/expected_since_2.1.0.toml | 7 +++- tools/differential/expected_since_2.2.0.toml | 7 +++- tools/differential/expected_since_2.3.0.toml | 7 +++- 8 files changed, 62 insertions(+), 12 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 25f2982d..4217b553 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -877,6 +877,9 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py): - 2026-09-28 (Derek), #544 — A PART READ WHOLLY AS SUFFIXES REPORTS NO NAME READING OF ITS WORDS. group's particle chain runs over every comma segment, and its two emitters — `particle-or-given` when a particle behind a word of both the title and the particle vocabulary chains (since 2.0.0, de264af1) and `suffix-or-name` when the chain takes an ambiguous acronym into the name (#289/#516, 59d8f38a, in no release) — reported inside a TAIL segment, which assign reads wholly as suffixes. Each such report named a reading the parse never made, against rules.md#A1's "A report names the reading the parse took". A family comma's tail was already silent, since group hands the chain no report list anywhere after a family comma; the suffix comma's tails were not. Measured on the released wheels: `John Smith, Jr., Freiherr von Richthofen` reports `particle-or-given` on 'von', a token in the suffix role, at 2.0.0, 2.1.0, 2.2.0 and 2.3.0; `John Smith, Jr., PhD van Ma` and `John Smith, Jr., PhD Do Ma` report it on 'van' and 'Do' at 2.0.0 and 2.1.0 only. The `suffix-or-name` half reached `John Smith, Jr., PhD Do Ma`, `John Smith, MA, PhD Do Ma` and `John Smith, Jr., PhD van Ma` on master (e10e83b4), and the run rule of the bullet above made it reachable behind ONE comma: `John Smith, PhD Do Ma` carried C1's flip and a second report on 'Ma'. FIXED by scope, not by a new test at the emitter: group passes the chain no report list in a tail segment either, so both emitters go quiet there together, and rules.md#C2 states the boundary for any part consumed wholly as suffixes. The maiden channel is a separate parameter and is untouched: a tail segment's reader is NONE, so the maiden walk reports nothing there to begin with. No mechanisms.md entry: this is AMBIGUITY-AT-THE-DECISION-SITE's own contract (a report fires only where the parse chose between live readings) applied to a stage whose reading a later stage overrides for the whole segment. MEASURED 2026-09-28, the tree against the same tree with `None if family_comma else ambiguities` restored in `group()` (the comparator), each parse recorded as its seven fields plus `(kind, [(token text, token role)])` per report, under all three name orders: 0 of the 1441 differential-corpus names move, so the gate has nothing to classify; over a comma grid — the prefixes `John Smith, `, `John Smith, Jr., ` and `Smith, John, ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, Do, van, de, Jr, MEng, Ed, y, i}, each text as written, lowercased and uppercased, deduplicated to 11,049 texts — 360 parses (120 texts, every one of them mixed case) lose one `suffix-or-name` report apiece, every removed report on a token in the suffix role, 0 reports added, 0 field moves. 21 of the 120 texts carry one comma and are the run rule's reach; the other 99 carry two and moved the same way (297 parses) when the same one-line change was applied to master e10e83b4; the `Smith, John, ` prefix moves nothing, being a family comma. The grid holds no word of both the title and the particle vocabulary, so the `particle-or-given` half is witnessed by the case row alone. Pinned by the case rows `a_credential_run_after_the_comma_reports_no_chain_fork` and `a_part_past_the_second_reports_no_particle_fork`; the rows the chain still reports on outside a tail are `the_chain_reports_the_acronym_it_takes` and `titled_particle_chain_survives_a_title_that_is_also_a_particle`. - 2026-10-01 (Derek), #562 — A PARTICLE CHAIN UNSETTLES A RUN THE CAPITALS SETTLED, AND THE COUNT READS IT. The 2026-09-27 bullet above left a run whose every member is listed and leans credential to the family-comma path, "which already reads it whole". That promise fails wherever two particles stand side by side in the part: group chains them into one particle run (P2), assign reads the part as name text, and P6 attaches the chain to the family — `John Smith, PhD DO DO` read given 'PhD', family 'DO DO John Smith', and `John Smith, DO DO DO` given 'DO', middle 'DO DO'. rules.md#S2 already said the capitals do not decide a member chained behind another particle ("the run attaches whatever the capitals say (P6)"), so C1's shortcut was resting on a premise S2 denies. Of the two fixes #562 weighed, the one taken narrows the shortcut and leaves S2 as written: a part holding two particles side by side is read by the count, which two name words before the comma flip to the credential run, reported (`suffix-or-name`). The other — letting C1's evidence or S2's credential-in-front company outrank the chain — would have contradicted S2's sentence and P6's `Doe, John van DO` example, so it needed S2 amended rather than a gap filled. The test is ANY two adjacent particles, not a member behind one: `vd` is a particle and an unambiguous suffix word, so `John Smith, PhD vd DO` and `John Smith, MA vd vd` chained and misread the same way, the second with no member behind a particle at all. Segment runs before classify, so it asks classify's own predicate (`_normalize(text) in lexicon.particles`) and only while the run is still settled. The test does not ask whether the family-comma path would actually have misread the part, which it could not without reading ahead to group: where that path did read the part whole — for example a pair opening the part with an unambiguous particle-and-suffix word (`John Smith, vd DO`, `John Smith, VD DO`), a credential that is also a title in front (`John Smith, MD DO DO`), or a credential closed by a period in front (`John Smith, Esq. DO DO`, `John Smith, Jr. DO DO`), the list being by example rather than a census — the count flips the part to the same fields and reports the call, as every flip at this comma does (rules.md#C1's "A decision either way at this comma is reported"). Those reports are ACCEPTED (Derek, 2026-10-01: none of these is a name anyone would write on purpose, so a report is the right signal): with the capitals no longer settling the run, the call is the count's, and the report says so; `tests/v2/cases.py` pins `John Smith, vd DO`. 1.4.0 read every one of these names as the fix does; 2.0.0 and 2.1.0 read `John Smith, PhD DO DO` as title 'PhD', given 'DO DO', and 2.2.0 and 2.3.0 as the issue describes. MEASURED 2026-10-01 against master 0eadedeb, py3.11, `nameparser.__file__` asserted on each side, each parse compared as its seven fields plus its sorted ambiguity kinds: 0 of the 1453 differential-corpus names move (the two names this change adds are the gate's only movers, under a new fix(#562) rule in the four 2.x ledgers); over tests/v2/test_properties.py's settled grid (5,580 texts) 15 move, 10 of them role moves, and the other 5 (`John Smith, MD DO DO`, `MS`, `Esq.`, `Sr`, `Ms` in front) keep their fields and gain the flip's report; over a wider grid — the prefixes `John Smith, `, `Smith, ` and `Doe, John ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, DO, Do, do, vd, van, Jr, MD, Ms}, 4,389 texts — 30 move, 17 of them role moves, every mover a `John Smith, ` text now reading given 'John', family 'Smith' and the whole part as suffix, and every one reporting `suffix-or-name`. Recompute: check out the parent into a separate worktree, parse each grid in both trees under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff. The settled grid's own pin moves with it: tests/v2/test_properties.py's `_SETTLED_COUNT` reads 2,994 where it read 3,024, the ten `_SETTLED_EXCEPTIONS` that pinned #562 are gone, and its two recorded negative controls read 1,142 and 215 — the second had already moved from 751 to 215 with #563, before this change. LEFT OPEN, as #562 asked: `Smith, PhD DO DO` (ONE name word before the comma, so the count keeps the listing form, and P6 attaches the chain: given 'PhD', family 'DO DO Smith'), where no rule states S2's credential-in-front company against a chain; and `John Smith, PhD van der`, whose particles are no members and never reach the run test, reading given 'PhD', family 'van der John Smith' as before. - 2026-10-01 (Derek) — A CAPITALS-SETTLED RUN OPENED BY A WORD OF THE CLASS REPORTS, AND RULES.MD#C1 NOW SAYS SO. C1 said a run whose every class word is written in capitals "reads whole in silence". That held only behind another credential: a class word OPENING the part is the first word after the comma, whose decision C1 reports either way, so `Doe, MA PhD`, `John Smith, MA MA` and `Smith, MA PhD` read wholly as suffixes and report `suffix-or-name` once, on that word, while `John Smith, PhD MA` and `Smith, PhD MA` are silent. The behavior predates #562 and is kept: nobody repeats `MA` at the end of their name on purpose, so the report is the right signal. Statement corrected, no parse moved. MEASURED 2026-10-01 on master b39c370c, the prefixes `John Smith, ` and `Smith, ` times every run of two or three words drawn with repetition from {MA, BA, ED, DO, JD, PhD, MD, Jr, Esq.} holding at least one class word: all 900 runs opening with a class word report, 895 of them read whole (14 of those hold a #562 particle pair, whose count flip reports over the whole part rather than on the first word) and the other five being `Smith, MA DO DO` and its like, the one-name-word chain #562 left open; of the 560 opening with another credential, 554 read whole in silence, and the other six hold a #562 particle pair — four reading whole and reporting (`John Smith, PhD DO DO`), two being `Smith, PhD DO DO` and `Smith, Jr DO DO`. Recompute: parse that grid and bucket by whether the first word is a class word, whether `ambiguities` is empty, and whether the suffix is the whole part. The same day (Derek) rules.md#S2's first-slot precedence sentence gained the pointer to that exception, so the next reader does not rediscover it: S2 says the count decides the first slot after a family comma before case, true for one word (`John Smith, MA` flips by the count) and false for a capitals-settled run, where C1's shortcut reads the capitals first and the comma keeps its family reading. And S2's list of reporting slots lost "the segments beyond it": no `suffix-or-name` report lands past the second comma (`Smith, John, MA` and `John Smith, Jr., MA` are silent). It once did: at #530's merge cc78c960 `John Smith, Jr., PhD Do Ma` reported `suffix-or-name` on 'Ma' beside its `comma-structure` flag, from group's chain emitter, which the 2026-09-28 #544 bullet above silenced in a tail. S2 now points at C2 for what such a part does report, which is not only `comma-structure` — an unclosed delimiter there adds `unbalanced-delimiter` (`John Smith, Jr., (Bob`), and a maiden clause standing in it is read as one (`Smith, John, Jr nee Jones MA`, maiden 'Jones MA'). Derek, 2026-10-01: the silence is right. +- 2026-10-01 (Derek), #554 — A WORD OF BOTH THE PARTICLE AND THE SUFFIX VOCABULARY COUNTS IN THE RUN, AND STILL SPEAKS FOR NOTHING. After #544 the run test (`_segment.py`) counts `vd` and `mc` as suffix words, while S2's company (`_pieces._anchors`) refuses to let them speak for the word behind them, an exclusion PR #552 added after `Smith vd Ma, John` lost the contiguity of its family name. The issue framed that as two answers to one question and offered (b): end the run at such a particle, so that one "has a live name reading" predicate decides both. DECIDED (a), the shipped behavior, and no parse moves. They are two questions. Run membership asks what the word IS, and `vd` is unambiguous suffix vocabulary; company asks whether the word SPEAKS FOR the one behind it, and a particle cannot, P2 joining it forward to exactly that word. Both answers follow from the word's vocabulary, and they differ by design rather than by drift; S2's "end the run" is the company's run, not the part C1 counts. The clause reaches exactly `vd` and `mc`: the third word in both vocabularies, `do`, is in the ambiguous half and so is already a word of C1's class, reporting where these two do not (`John Smith, Jr do` against `John Smith, Jr vd`). Recompute the set: `L = Parser().lexicon; L.particles & (L.suffix_acronyms | L.suffix_words) - L.suffix_acronyms_ambiguous`. + The evidence is the writing that carries no other signal. Neither reading of `vd Ma` after a comma is realistic, a Dutch particle in front of a Chinese family name being as rare as `vd` or `mc` as a credential, so the mixed-case writing is an arbitrary edge case and decides nothing. Written in one case, where S2 has no capitals to read and the vocabulary is all there is, the part reads as the credential run, and that is the reading a reader would take: measured 2026-10-01 on master (fc682e36), `JOHN SMITH, VD MA` and `john smith, vd ma` read given 'JOHN'/'john', family 'SMITH'/'smith', suffix 'VD MA'/'vd ma' and report `suffix-or-name`, as `John Smith, vd Ma` does, and `John Smith, mc Ma` reads the same in every casing. (b) would add a particle exception to C1, and to #563's paired-initials test besides, in order to give `JOHN SMITH, VD MA` family 'JOHN SMITH VD MA' with no given name: the worse reading, on exactly the input where nothing else speaks. + #563's site needs nothing of its own. `García Márquez, vd G.J.` reads as `García Márquez, PhD G.J.` does, in every casing: given 'García', family 'Márquez', suffix 'vd G.J.', and neither reports. If a surname split by a credential run in silence is a defect, it is the run's, not the particle's. Recompute: `parse(s)` with `as_dict()` and `[a.kind.value for a in parse(s).ambiguities]` over each string above and its `.upper()` and `.lower()`. ### T1 — separators, not joiners diff --git a/docs/design/rules.md b/docs/design/rules.md index 683fda0e..d400e464 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1704,6 +1704,11 @@ C1. Rationale: a credential run after the comma means the name is in and S3 retires single-character matches for the same reason, while a multi-letter numeral or generational word stands in a run like any suffix word ('John Smith, III Ma'). A word of both + the particle and the unambiguous suffix vocabulary counts as a + suffix word there too ('John Smith, vd Ma'): membership in this + part asks what the word is, while S2's company asks whether it + speaks for the word behind it, which a particle never does, so + the company run it ends is S2's and not this one. A word of both the title and the suffix vocabulary opening such a part counts as a suffix word there, the name before the comma being complete, except as the titles of paired initials, above. Where every word @@ -1812,6 +1817,8 @@ C1. Rationale: a credential run after the comma means the name is in "John Smith, PhD DO DO" → suffix="PhD DO DO" "John Smith, PhD DO DO" → ambiguities=("suffix-or-name",) "John Smith, PhD vd DO" → suffix="PhD vd DO" + "John Smith, vd Ma" → suffix="vd Ma" + "JOHN SMITH, VD MA" → suffix="VD MA" "John Smith, X.Y.Z." → suffix="X.Y.Z." "John Smith, X.Y.Z." → ambiguities=() "John Smith, X.Y.Z." unlisted_dotted_suffixes-off → given="X.Y.Z." diff --git a/tests/v2/test_ledger_guards.py b/tests/v2/test_ledger_guards.py index 909d6496..93b66389 100644 --- a/tests/v2/test_ledger_guards.py +++ b/tests/v2/test_ledger_guards.py @@ -3067,8 +3067,12 @@ class _LatinCopy(NamedTuple): frozenset({"Doe, Jane nee Smith PhD MEng", "Jane Doe nee Smith PhD MA", "Jane Doe nee Smith PhD MEng"}), frozenset({r"jack\s+m\.a\.", r"wang\s+m\.eng\."}), - frozenset({"John Smith, Ed Ma", "John Smith, Ms Ma", - "John Smith, PhD Ma", r"John Smith, X\.Y\.Z\. MA", + # 2026-10-01, #554: the run rule gains rules.md#C1's particle-and- + # suffix examples, a vocabulary word in front of a member, which + # copies no set either. + frozenset({"JOHN SMITH, VD MA", "John Smith, Ed Ma", + "John Smith, Ms Ma", "John Smith, PhD Ma", + r"John Smith, X\.Y\.Z\. MA", "John Smith, vd Ma", "john smith, md ma"}), frozenset({"Doe, Jane PhD MEng", "Doe, Jane nee Smith PhD MEng", "Jane Doe nee Smith PhD MEng", "John Smith PhD MEng"}), @@ -3761,8 +3765,11 @@ def _claim(rule: dict) -> _Claim: # 2026-10-01, #562: 408 -> 410, 'John Smith, PhD DO DO' # and 'John Smith, PhD vd DO', #562's rules.md#C1 examples # and case rows. Reach, verified name by name. + # 2026-10-01, #554: 410 -> 412, 'John Smith, vd Ma' and + # 'JOHN SMITH, VD MA', #554's rules.md#C1 examples. Reach, + # verified name by name. "fix(comma-family) lone post-comma piece routes to suffix/title, not first": - _Claim(410, ('given', 'suffix', 'title'), "2b67e34c920c", None), + _Claim(412, ('given', 'suffix', 'title'), "d9217453f0df", None), "fix(comma-family) a comma followed only by titles keeps the given/family split": _Claim(2, ('family', 'given'), "5bd9c6d96c38", None), "fix(comma-family) a comma followed only by titles keeps the given/family split, the C1 example": @@ -3857,8 +3864,11 @@ def _claim(rule: dict) -> _Claim: # 2026-10-01, #562: 408 -> 410, 'John Smith, PhD DO DO' # and 'John Smith, PhD vd DO', #562's rules.md#C1 examples # and case rows. Reach, verified name by name. + # 2026-10-01, #554: 410 -> 412, 'John Smith, vd Ma' and + # 'JOHN SMITH, VD MA', #554's rules.md#C1 examples. Reach, + # verified name by name. "fix(comma-precomma-family) pre-comma run reads as family, not given": - _Claim(410, ('family', 'given'), "2b67e34c920c", None), + _Claim(412, ('family', 'given'), "d9217453f0df", None), # 2026-09-20, #397: retitled in place, reach and digest # unchanged -- the rule keeps 'Carod i', which the landing # leaves byte-identical. @@ -5190,8 +5200,10 @@ def _claim(rule: dict) -> _Claim: # 2026-09-27, #544: new, 4; gains 'John Smith, Ed Ma', 'John # Smith, Ms Ma', 'John Smith, X.Y.Z. MA', 'john smith, md ma'. # 2026-09-28, #544: 4 -> 5; gains 'John Smith, PhD Ma'. + # 2026-10-01, #554: 5 -> 7; gains 'John Smith, vd Ma' and + # 'JOHN SMITH, VD MA'. "fix(#544) a run of ambiguous members after a suffix comma reads by the name-word count": - _Claim(5, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "4123358beccd", ('DEFAULT',)), + _Claim(7, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "890e7723c601", ('DEFAULT',)), # 2026-09-27, #544: new, 4; gains 'Doe, Jane PhD MEng', 'Doe, # Jane nee Smith PhD MEng', 'Jane Doe nee Smith PhD MEng', # 'John Smith PhD MEng'. Relabelled the same day @@ -5640,8 +5652,10 @@ def _claim(rule: dict) -> _Claim: # 2026-09-27, #544: new, 4; gains 'John Smith, Ed Ma', 'John # Smith, Ms Ma', 'John Smith, X.Y.Z. MA', 'john smith, md ma'. # 2026-09-28, #544: 4 -> 5; gains 'John Smith, PhD Ma'. + # 2026-10-01, #554: 5 -> 7; gains 'John Smith, vd Ma' and + # 'JOHN SMITH, VD MA'. "fix(#544) a run of ambiguous members after a suffix comma reads by the name-word count": - _Claim(5, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "4123358beccd", ('DEFAULT',)), + _Claim(7, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "890e7723c601", ('DEFAULT',)), # 2026-09-27, #544: new, 4; gains 'Doe, Jane PhD MEng', 'Doe, # Jane nee Smith PhD MEng', 'Jane Doe nee Smith PhD MEng', # 'John Smith PhD MEng'. Relabelled the same day @@ -6324,8 +6338,10 @@ def _claim(rule: dict) -> _Claim: # 2026-09-27, #544: new, 4; gains 'John Smith, Ed Ma', 'John # Smith, Ms Ma', 'John Smith, X.Y.Z. MA', 'john smith, md ma'. # 2026-09-28, #544: 4 -> 5; gains 'John Smith, PhD Ma'. + # 2026-10-01, #554: 5 -> 7; gains 'John Smith, vd Ma' and + # 'JOHN SMITH, VD MA'. "fix(#544) a run of ambiguous members after a suffix comma reads by the name-word count": - _Claim(5, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "4123358beccd", ('DEFAULT',)), + _Claim(7, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "890e7723c601", ('DEFAULT',)), # 2026-09-27, #544: new, 4; gains 'Doe, Jane PhD MEng', 'Doe, # Jane nee Smith PhD MEng', 'Jane Doe nee Smith PhD MEng', # 'John Smith PhD MEng'. Relabelled the same day @@ -6623,8 +6639,10 @@ def _claim(rule: dict) -> _Claim: # 2026-09-27, #544: new, 4; gains 'John Smith, Ed Ma', 'John # Smith, Ms Ma', 'John Smith, X.Y.Z. MA', 'john smith, md ma'. # 2026-09-28, #544: 4 -> 5; gains 'John Smith, PhD Ma'. + # 2026-10-01, #554: 5 -> 7; gains 'John Smith, vd Ma' and + # 'JOHN SMITH, VD MA'. "fix(#544) a run of ambiguous members after a suffix comma reads by the name-word count": - _Claim(5, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "4123358beccd", ('DEFAULT',)), + _Claim(7, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "890e7723c601", ('DEFAULT',)), # 2026-09-27, #544: new, 4; gains 'Doe, Jane PhD MEng', 'Doe, # Jane nee Smith PhD MEng', 'Jane Doe nee Smith PhD MEng', # 'John Smith PhD MEng'. diff --git a/tools/differential/corpus_rules.jsonl b/tools/differential/corpus_rules.jsonl index 5701e1b6..0d266b23 100644 --- a/tools/differential/corpus_rules.jsonl +++ b/tools/differential/corpus_rules.jsonl @@ -109,6 +109,7 @@ "JOHN QUINCY SMITH I" "JOHN SMITH MA" "JOHN SMITH PH.D." +"JOHN SMITH, VD MA" "JOSE ORTEGA-Y-GASSET" "JUAN GARCIA Jr." "JUAN GARCIA Y LOPEZ" @@ -223,6 +224,7 @@ "John Smith, V." "John Smith, X.Y. P.Q." "John Smith, X.Y.Z." +"John Smith, vd Ma" "John and Jane Smith" "John née Jones Smith MA" "John née Jones Smith Ma" diff --git a/tools/differential/expected_since_2.0.0.toml b/tools/differential/expected_since_2.0.0.toml index 3af76509..61a39228 100644 --- a/tools/differential/expected_since_2.0.0.toml +++ b/tools/differential/expected_since_2.0.0.toml @@ -3768,11 +3768,16 @@ issue = "fix(#544) a run of ambiguous members after a suffix comma reads by the # unambiguous word in front of the member, which the count reads whole # like the rest. # +# 2026-10-01, #554: 'John Smith, vd Ma' and 'JOHN SMITH, VD MA' join, +# rules.md#C1's examples for a word of both the particle and the suffix +# vocabulary, which "counts as a suffix word there too". This baseline +# read the whole name as the family ('John Smith vd Ma'). +# # Literal; `fields` is the union the names move ('title' is the duals'). # Probes: 'Smith, Ed Ma' (one name word before the comma keeps the # listing form), 'Smith, Ms Ma' (a dual opening the given part is a # title) and the superstring 'Dr. John Smith, Ed Ma' are _MUST_NOT_MATCH. -name_regex = "^(?:John Smith, Ed Ma|John Smith, Ms Ma|John Smith, PhD Ma|John Smith, X\\.Y\\.Z\\. MA|john smith, md ma)$" +name_regex = "^(?:JOHN SMITH, VD MA|John Smith, Ed Ma|John Smith, Ms Ma|John Smith, PhD Ma|John Smith, X\\.Y\\.Z\\. MA|John Smith, vd Ma|john smith, md ma)$" fields = ["given", "middle", "family", "suffix", "title", "_ambiguities"] orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.1.0.toml b/tools/differential/expected_since_2.1.0.toml index 0d9b6c1a..1e624fe9 100644 --- a/tools/differential/expected_since_2.1.0.toml +++ b/tools/differential/expected_since_2.1.0.toml @@ -3679,11 +3679,16 @@ issue = "fix(#544) a run of ambiguous members after a suffix comma reads by the # unambiguous word in front of the member, which the count reads whole # like the rest. # +# 2026-10-01, #554: 'John Smith, vd Ma' and 'JOHN SMITH, VD MA' join, +# rules.md#C1's examples for a word of both the particle and the suffix +# vocabulary, which "counts as a suffix word there too". This baseline +# read the whole name as the family ('John Smith vd Ma'). +# # Literal; `fields` is the union the names move ('title' is the duals'). # Probes: 'Smith, Ed Ma' (one name word before the comma keeps the # listing form), 'Smith, Ms Ma' (a dual opening the given part is a # title) and the superstring 'Dr. John Smith, Ed Ma' are _MUST_NOT_MATCH. -name_regex = "^(?:John Smith, Ed Ma|John Smith, Ms Ma|John Smith, PhD Ma|John Smith, X\\.Y\\.Z\\. MA|john smith, md ma)$" +name_regex = "^(?:JOHN SMITH, VD MA|John Smith, Ed Ma|John Smith, Ms Ma|John Smith, PhD Ma|John Smith, X\\.Y\\.Z\\. MA|John Smith, vd Ma|john smith, md ma)$" fields = ["given", "middle", "family", "suffix", "title", "_ambiguities"] orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.2.0.toml b/tools/differential/expected_since_2.2.0.toml index f75d0b52..784ff85a 100644 --- a/tools/differential/expected_since_2.2.0.toml +++ b/tools/differential/expected_since_2.2.0.toml @@ -2077,11 +2077,16 @@ issue = "fix(#544) a run of ambiguous members after a suffix comma reads by the # unambiguous word in front of the member, which the count reads whole # like the rest. # +# 2026-10-01, #554: 'John Smith, vd Ma' and 'JOHN SMITH, VD MA' join, +# rules.md#C1's examples for a word of both the particle and the suffix +# vocabulary, which "counts as a suffix word there too". This baseline +# read the whole name as the family ('John Smith vd Ma'). +# # Literal; `fields` is the union the names move ('title' is the duals'). # Probes: 'Smith, Ed Ma' (one name word before the comma keeps the # listing form), 'Smith, Ms Ma' (a dual opening the given part is a # title) and the superstring 'Dr. John Smith, Ed Ma' are _MUST_NOT_MATCH. -name_regex = "^(?:John Smith, Ed Ma|John Smith, Ms Ma|John Smith, PhD Ma|John Smith, X\\.Y\\.Z\\. MA|john smith, md ma)$" +name_regex = "^(?:JOHN SMITH, VD MA|John Smith, Ed Ma|John Smith, Ms Ma|John Smith, PhD Ma|John Smith, X\\.Y\\.Z\\. MA|John Smith, vd Ma|john smith, md ma)$" fields = ["given", "middle", "family", "suffix", "title", "_ambiguities"] orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.3.0.toml b/tools/differential/expected_since_2.3.0.toml index 89f5b803..99ee3e58 100644 --- a/tools/differential/expected_since_2.3.0.toml +++ b/tools/differential/expected_since_2.3.0.toml @@ -1344,11 +1344,16 @@ issue = "fix(#544) a run of ambiguous members after a suffix comma reads by the # unambiguous word in front of the member, which the count reads whole # like the rest. # +# 2026-10-01, #554: 'John Smith, vd Ma' and 'JOHN SMITH, VD MA' join, +# rules.md#C1's examples for a word of both the particle and the suffix +# vocabulary, which "counts as a suffix word there too". This baseline +# read the whole name as the family ('John Smith vd Ma'). +# # Literal; `fields` is the union the names move ('title' is the duals'). # Probes: 'Smith, Ed Ma' (one name word before the comma keeps the # listing form), 'Smith, Ms Ma' (a dual opening the given part is a # title) and the superstring 'Dr. John Smith, Ed Ma' are _MUST_NOT_MATCH. -name_regex = "^(?:John Smith, Ed Ma|John Smith, Ms Ma|John Smith, PhD Ma|John Smith, X\\.Y\\.Z\\. MA|john smith, md ma)$" +name_regex = "^(?:JOHN SMITH, VD MA|John Smith, Ed Ma|John Smith, Ms Ma|John Smith, PhD Ma|John Smith, X\\.Y\\.Z\\. MA|John Smith, vd Ma|john smith, md ma)$" fields = ["given", "middle", "family", "suffix", "title", "_ambiguities"] orders = ["DEFAULT"] From 3ad584ec12c321379955b2d38110602e2c3f1897 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 15:23:05 -0700 Subject: [PATCH 2/4] docs(C1): point #554's decision at follow-up #573 Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 4217b553..ee67c0c6 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -879,7 +879,7 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py): - 2026-10-01 (Derek) — A CAPITALS-SETTLED RUN OPENED BY A WORD OF THE CLASS REPORTS, AND RULES.MD#C1 NOW SAYS SO. C1 said a run whose every class word is written in capitals "reads whole in silence". That held only behind another credential: a class word OPENING the part is the first word after the comma, whose decision C1 reports either way, so `Doe, MA PhD`, `John Smith, MA MA` and `Smith, MA PhD` read wholly as suffixes and report `suffix-or-name` once, on that word, while `John Smith, PhD MA` and `Smith, PhD MA` are silent. The behavior predates #562 and is kept: nobody repeats `MA` at the end of their name on purpose, so the report is the right signal. Statement corrected, no parse moved. MEASURED 2026-10-01 on master b39c370c, the prefixes `John Smith, ` and `Smith, ` times every run of two or three words drawn with repetition from {MA, BA, ED, DO, JD, PhD, MD, Jr, Esq.} holding at least one class word: all 900 runs opening with a class word report, 895 of them read whole (14 of those hold a #562 particle pair, whose count flip reports over the whole part rather than on the first word) and the other five being `Smith, MA DO DO` and its like, the one-name-word chain #562 left open; of the 560 opening with another credential, 554 read whole in silence, and the other six hold a #562 particle pair — four reading whole and reporting (`John Smith, PhD DO DO`), two being `Smith, PhD DO DO` and `Smith, Jr DO DO`. Recompute: parse that grid and bucket by whether the first word is a class word, whether `ambiguities` is empty, and whether the suffix is the whole part. The same day (Derek) rules.md#S2's first-slot precedence sentence gained the pointer to that exception, so the next reader does not rediscover it: S2 says the count decides the first slot after a family comma before case, true for one word (`John Smith, MA` flips by the count) and false for a capitals-settled run, where C1's shortcut reads the capitals first and the comma keeps its family reading. And S2's list of reporting slots lost "the segments beyond it": no `suffix-or-name` report lands past the second comma (`Smith, John, MA` and `John Smith, Jr., MA` are silent). It once did: at #530's merge cc78c960 `John Smith, Jr., PhD Do Ma` reported `suffix-or-name` on 'Ma' beside its `comma-structure` flag, from group's chain emitter, which the 2026-09-28 #544 bullet above silenced in a tail. S2 now points at C2 for what such a part does report, which is not only `comma-structure` — an unclosed delimiter there adds `unbalanced-delimiter` (`John Smith, Jr., (Bob`), and a maiden clause standing in it is read as one (`Smith, John, Jr nee Jones MA`, maiden 'Jones MA'). Derek, 2026-10-01: the silence is right. - 2026-10-01 (Derek), #554 — A WORD OF BOTH THE PARTICLE AND THE SUFFIX VOCABULARY COUNTS IN THE RUN, AND STILL SPEAKS FOR NOTHING. After #544 the run test (`_segment.py`) counts `vd` and `mc` as suffix words, while S2's company (`_pieces._anchors`) refuses to let them speak for the word behind them, an exclusion PR #552 added after `Smith vd Ma, John` lost the contiguity of its family name. The issue framed that as two answers to one question and offered (b): end the run at such a particle, so that one "has a live name reading" predicate decides both. DECIDED (a), the shipped behavior, and no parse moves. They are two questions. Run membership asks what the word IS, and `vd` is unambiguous suffix vocabulary; company asks whether the word SPEAKS FOR the one behind it, and a particle cannot, P2 joining it forward to exactly that word. Both answers follow from the word's vocabulary, and they differ by design rather than by drift; S2's "end the run" is the company's run, not the part C1 counts. The clause reaches exactly `vd` and `mc`: the third word in both vocabularies, `do`, is in the ambiguous half and so is already a word of C1's class, reporting where these two do not (`John Smith, Jr do` against `John Smith, Jr vd`). Recompute the set: `L = Parser().lexicon; L.particles & (L.suffix_acronyms | L.suffix_words) - L.suffix_acronyms_ambiguous`. The evidence is the writing that carries no other signal. Neither reading of `vd Ma` after a comma is realistic, a Dutch particle in front of a Chinese family name being as rare as `vd` or `mc` as a credential, so the mixed-case writing is an arbitrary edge case and decides nothing. Written in one case, where S2 has no capitals to read and the vocabulary is all there is, the part reads as the credential run, and that is the reading a reader would take: measured 2026-10-01 on master (fc682e36), `JOHN SMITH, VD MA` and `john smith, vd ma` read given 'JOHN'/'john', family 'SMITH'/'smith', suffix 'VD MA'/'vd ma' and report `suffix-or-name`, as `John Smith, vd Ma` does, and `John Smith, mc Ma` reads the same in every casing. (b) would add a particle exception to C1, and to #563's paired-initials test besides, in order to give `JOHN SMITH, VD MA` family 'JOHN SMITH VD MA' with no given name: the worse reading, on exactly the input where nothing else speaks. - #563's site needs nothing of its own. `García Márquez, vd G.J.` reads as `García Márquez, PhD G.J.` does, in every casing: given 'García', family 'Márquez', suffix 'vd G.J.', and neither reports. If a surname split by a credential run in silence is a defect, it is the run's, not the particle's. Recompute: `parse(s)` with `as_dict()` and `[a.kind.value for a in parse(s).ambiguities]` over each string above and its `.upper()` and `.lower()`. + #563's site needs nothing of its own. `García Márquez, vd G.J.` reads as `García Márquez, PhD G.J.` does, in every casing: given 'García', family 'Márquez', suffix 'vd G.J.', and neither reports. If a surname split by a credential run in silence is a defect, it is the run's, not the particle's. Open: #573 — the same words in uniform case before the comma and in the given part's trailing slot, plus the issue's aside (`Doe, Jane PhD vd Ma` silent in middle). Recompute: `parse(s)` with `as_dict()` and `[a.kind.value for a in parse(s).ambiguities]` over each string above and its `.upper()` and `.lower()`. ### T1 — separators, not joiners From 06595241877f832d1d7045eee56de165e533d035 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 18:10:23 -0700 Subject: [PATCH 3/4] docs(C1): review round for #554 -- clearer clause, pinned report and scope - rules.md#C1: the clause now leads with what it decides (a word's particle reading does not take it out of the run) and names the two runs plainly; "such a part" gets an explicit antecedent. - New C1 examples pin the uniform-case report and the clause's scope: 'John Smith, Jr vd' is silent, 'John Smith, Jr do' reports (do is a word of the ambiguous class, so the clause does not reach it). - decisions.md#C1: PR #552's split was an intermediate state, never released; 1.4.0 read these names as (a) does (verified against the wheel); do's membership shown by a recompute; shorter Open: handle. - Ledgers: comments carry the "unambiguous" qualifier and quote the new clause; 'John Smith, Jr do' joins the #544 run rule at 2.0.0-2.3.0, with each baseline's own reading. Moved corpus claims re-recorded. Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 6 +- docs/design/rules.md | 23 ++++--- tests/v2/test_ledger_guards.py | 72 ++++++++++++-------- tools/differential/corpus_rules.jsonl | 2 + tools/differential/expected_since_2.0.0.toml | 12 ++-- tools/differential/expected_since_2.1.0.toml | 12 ++-- tools/differential/expected_since_2.2.0.toml | 12 ++-- tools/differential/expected_since_2.3.0.toml | 12 ++-- 8 files changed, 96 insertions(+), 55 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index ee67c0c6..62001bc4 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -877,9 +877,9 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py): - 2026-09-28 (Derek), #544 — A PART READ WHOLLY AS SUFFIXES REPORTS NO NAME READING OF ITS WORDS. group's particle chain runs over every comma segment, and its two emitters — `particle-or-given` when a particle behind a word of both the title and the particle vocabulary chains (since 2.0.0, de264af1) and `suffix-or-name` when the chain takes an ambiguous acronym into the name (#289/#516, 59d8f38a, in no release) — reported inside a TAIL segment, which assign reads wholly as suffixes. Each such report named a reading the parse never made, against rules.md#A1's "A report names the reading the parse took". A family comma's tail was already silent, since group hands the chain no report list anywhere after a family comma; the suffix comma's tails were not. Measured on the released wheels: `John Smith, Jr., Freiherr von Richthofen` reports `particle-or-given` on 'von', a token in the suffix role, at 2.0.0, 2.1.0, 2.2.0 and 2.3.0; `John Smith, Jr., PhD van Ma` and `John Smith, Jr., PhD Do Ma` report it on 'van' and 'Do' at 2.0.0 and 2.1.0 only. The `suffix-or-name` half reached `John Smith, Jr., PhD Do Ma`, `John Smith, MA, PhD Do Ma` and `John Smith, Jr., PhD van Ma` on master (e10e83b4), and the run rule of the bullet above made it reachable behind ONE comma: `John Smith, PhD Do Ma` carried C1's flip and a second report on 'Ma'. FIXED by scope, not by a new test at the emitter: group passes the chain no report list in a tail segment either, so both emitters go quiet there together, and rules.md#C2 states the boundary for any part consumed wholly as suffixes. The maiden channel is a separate parameter and is untouched: a tail segment's reader is NONE, so the maiden walk reports nothing there to begin with. No mechanisms.md entry: this is AMBIGUITY-AT-THE-DECISION-SITE's own contract (a report fires only where the parse chose between live readings) applied to a stage whose reading a later stage overrides for the whole segment. MEASURED 2026-09-28, the tree against the same tree with `None if family_comma else ambiguities` restored in `group()` (the comparator), each parse recorded as its seven fields plus `(kind, [(token text, token role)])` per report, under all three name orders: 0 of the 1441 differential-corpus names move, so the gate has nothing to classify; over a comma grid — the prefixes `John Smith, `, `John Smith, Jr., ` and `Smith, John, ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, Do, van, de, Jr, MEng, Ed, y, i}, each text as written, lowercased and uppercased, deduplicated to 11,049 texts — 360 parses (120 texts, every one of them mixed case) lose one `suffix-or-name` report apiece, every removed report on a token in the suffix role, 0 reports added, 0 field moves. 21 of the 120 texts carry one comma and are the run rule's reach; the other 99 carry two and moved the same way (297 parses) when the same one-line change was applied to master e10e83b4; the `Smith, John, ` prefix moves nothing, being a family comma. The grid holds no word of both the title and the particle vocabulary, so the `particle-or-given` half is witnessed by the case row alone. Pinned by the case rows `a_credential_run_after_the_comma_reports_no_chain_fork` and `a_part_past_the_second_reports_no_particle_fork`; the rows the chain still reports on outside a tail are `the_chain_reports_the_acronym_it_takes` and `titled_particle_chain_survives_a_title_that_is_also_a_particle`. - 2026-10-01 (Derek), #562 — A PARTICLE CHAIN UNSETTLES A RUN THE CAPITALS SETTLED, AND THE COUNT READS IT. The 2026-09-27 bullet above left a run whose every member is listed and leans credential to the family-comma path, "which already reads it whole". That promise fails wherever two particles stand side by side in the part: group chains them into one particle run (P2), assign reads the part as name text, and P6 attaches the chain to the family — `John Smith, PhD DO DO` read given 'PhD', family 'DO DO John Smith', and `John Smith, DO DO DO` given 'DO', middle 'DO DO'. rules.md#S2 already said the capitals do not decide a member chained behind another particle ("the run attaches whatever the capitals say (P6)"), so C1's shortcut was resting on a premise S2 denies. Of the two fixes #562 weighed, the one taken narrows the shortcut and leaves S2 as written: a part holding two particles side by side is read by the count, which two name words before the comma flip to the credential run, reported (`suffix-or-name`). The other — letting C1's evidence or S2's credential-in-front company outrank the chain — would have contradicted S2's sentence and P6's `Doe, John van DO` example, so it needed S2 amended rather than a gap filled. The test is ANY two adjacent particles, not a member behind one: `vd` is a particle and an unambiguous suffix word, so `John Smith, PhD vd DO` and `John Smith, MA vd vd` chained and misread the same way, the second with no member behind a particle at all. Segment runs before classify, so it asks classify's own predicate (`_normalize(text) in lexicon.particles`) and only while the run is still settled. The test does not ask whether the family-comma path would actually have misread the part, which it could not without reading ahead to group: where that path did read the part whole — for example a pair opening the part with an unambiguous particle-and-suffix word (`John Smith, vd DO`, `John Smith, VD DO`), a credential that is also a title in front (`John Smith, MD DO DO`), or a credential closed by a period in front (`John Smith, Esq. DO DO`, `John Smith, Jr. DO DO`), the list being by example rather than a census — the count flips the part to the same fields and reports the call, as every flip at this comma does (rules.md#C1's "A decision either way at this comma is reported"). Those reports are ACCEPTED (Derek, 2026-10-01: none of these is a name anyone would write on purpose, so a report is the right signal): with the capitals no longer settling the run, the call is the count's, and the report says so; `tests/v2/cases.py` pins `John Smith, vd DO`. 1.4.0 read every one of these names as the fix does; 2.0.0 and 2.1.0 read `John Smith, PhD DO DO` as title 'PhD', given 'DO DO', and 2.2.0 and 2.3.0 as the issue describes. MEASURED 2026-10-01 against master 0eadedeb, py3.11, `nameparser.__file__` asserted on each side, each parse compared as its seven fields plus its sorted ambiguity kinds: 0 of the 1453 differential-corpus names move (the two names this change adds are the gate's only movers, under a new fix(#562) rule in the four 2.x ledgers); over tests/v2/test_properties.py's settled grid (5,580 texts) 15 move, 10 of them role moves, and the other 5 (`John Smith, MD DO DO`, `MS`, `Esq.`, `Sr`, `Ms` in front) keep their fields and gain the flip's report; over a wider grid — the prefixes `John Smith, `, `Smith, ` and `Doe, John ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, DO, Do, do, vd, van, Jr, MD, Ms}, 4,389 texts — 30 move, 17 of them role moves, every mover a `John Smith, ` text now reading given 'John', family 'Smith' and the whole part as suffix, and every one reporting `suffix-or-name`. Recompute: check out the parent into a separate worktree, parse each grid in both trees under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff. The settled grid's own pin moves with it: tests/v2/test_properties.py's `_SETTLED_COUNT` reads 2,994 where it read 3,024, the ten `_SETTLED_EXCEPTIONS` that pinned #562 are gone, and its two recorded negative controls read 1,142 and 215 — the second had already moved from 751 to 215 with #563, before this change. LEFT OPEN, as #562 asked: `Smith, PhD DO DO` (ONE name word before the comma, so the count keeps the listing form, and P6 attaches the chain: given 'PhD', family 'DO DO Smith'), where no rule states S2's credential-in-front company against a chain; and `John Smith, PhD van der`, whose particles are no members and never reach the run test, reading given 'PhD', family 'van der John Smith' as before. - 2026-10-01 (Derek) — A CAPITALS-SETTLED RUN OPENED BY A WORD OF THE CLASS REPORTS, AND RULES.MD#C1 NOW SAYS SO. C1 said a run whose every class word is written in capitals "reads whole in silence". That held only behind another credential: a class word OPENING the part is the first word after the comma, whose decision C1 reports either way, so `Doe, MA PhD`, `John Smith, MA MA` and `Smith, MA PhD` read wholly as suffixes and report `suffix-or-name` once, on that word, while `John Smith, PhD MA` and `Smith, PhD MA` are silent. The behavior predates #562 and is kept: nobody repeats `MA` at the end of their name on purpose, so the report is the right signal. Statement corrected, no parse moved. MEASURED 2026-10-01 on master b39c370c, the prefixes `John Smith, ` and `Smith, ` times every run of two or three words drawn with repetition from {MA, BA, ED, DO, JD, PhD, MD, Jr, Esq.} holding at least one class word: all 900 runs opening with a class word report, 895 of them read whole (14 of those hold a #562 particle pair, whose count flip reports over the whole part rather than on the first word) and the other five being `Smith, MA DO DO` and its like, the one-name-word chain #562 left open; of the 560 opening with another credential, 554 read whole in silence, and the other six hold a #562 particle pair — four reading whole and reporting (`John Smith, PhD DO DO`), two being `Smith, PhD DO DO` and `Smith, Jr DO DO`. Recompute: parse that grid and bucket by whether the first word is a class word, whether `ambiguities` is empty, and whether the suffix is the whole part. The same day (Derek) rules.md#S2's first-slot precedence sentence gained the pointer to that exception, so the next reader does not rediscover it: S2 says the count decides the first slot after a family comma before case, true for one word (`John Smith, MA` flips by the count) and false for a capitals-settled run, where C1's shortcut reads the capitals first and the comma keeps its family reading. And S2's list of reporting slots lost "the segments beyond it": no `suffix-or-name` report lands past the second comma (`Smith, John, MA` and `John Smith, Jr., MA` are silent). It once did: at #530's merge cc78c960 `John Smith, Jr., PhD Do Ma` reported `suffix-or-name` on 'Ma' beside its `comma-structure` flag, from group's chain emitter, which the 2026-09-28 #544 bullet above silenced in a tail. S2 now points at C2 for what such a part does report, which is not only `comma-structure` — an unclosed delimiter there adds `unbalanced-delimiter` (`John Smith, Jr., (Bob`), and a maiden clause standing in it is read as one (`Smith, John, Jr nee Jones MA`, maiden 'Jones MA'). Derek, 2026-10-01: the silence is right. -- 2026-10-01 (Derek), #554 — A WORD OF BOTH THE PARTICLE AND THE SUFFIX VOCABULARY COUNTS IN THE RUN, AND STILL SPEAKS FOR NOTHING. After #544 the run test (`_segment.py`) counts `vd` and `mc` as suffix words, while S2's company (`_pieces._anchors`) refuses to let them speak for the word behind them, an exclusion PR #552 added after `Smith vd Ma, John` lost the contiguity of its family name. The issue framed that as two answers to one question and offered (b): end the run at such a particle, so that one "has a live name reading" predicate decides both. DECIDED (a), the shipped behavior, and no parse moves. They are two questions. Run membership asks what the word IS, and `vd` is unambiguous suffix vocabulary; company asks whether the word SPEAKS FOR the one behind it, and a particle cannot, P2 joining it forward to exactly that word. Both answers follow from the word's vocabulary, and they differ by design rather than by drift; S2's "end the run" is the company's run, not the part C1 counts. The clause reaches exactly `vd` and `mc`: the third word in both vocabularies, `do`, is in the ambiguous half and so is already a word of C1's class, reporting where these two do not (`John Smith, Jr do` against `John Smith, Jr vd`). Recompute the set: `L = Parser().lexicon; L.particles & (L.suffix_acronyms | L.suffix_words) - L.suffix_acronyms_ambiguous`. - The evidence is the writing that carries no other signal. Neither reading of `vd Ma` after a comma is realistic, a Dutch particle in front of a Chinese family name being as rare as `vd` or `mc` as a credential, so the mixed-case writing is an arbitrary edge case and decides nothing. Written in one case, where S2 has no capitals to read and the vocabulary is all there is, the part reads as the credential run, and that is the reading a reader would take: measured 2026-10-01 on master (fc682e36), `JOHN SMITH, VD MA` and `john smith, vd ma` read given 'JOHN'/'john', family 'SMITH'/'smith', suffix 'VD MA'/'vd ma' and report `suffix-or-name`, as `John Smith, vd Ma` does, and `John Smith, mc Ma` reads the same in every casing. (b) would add a particle exception to C1, and to #563's paired-initials test besides, in order to give `JOHN SMITH, VD MA` family 'JOHN SMITH VD MA' with no given name: the worse reading, on exactly the input where nothing else speaks. - #563's site needs nothing of its own. `García Márquez, vd G.J.` reads as `García Márquez, PhD G.J.` does, in every casing: given 'García', family 'Márquez', suffix 'vd G.J.', and neither reports. If a surname split by a credential run in silence is a defect, it is the run's, not the particle's. Open: #573 — the same words in uniform case before the comma and in the given part's trailing slot, plus the issue's aside (`Doe, Jane PhD vd Ma` silent in middle). Recompute: `parse(s)` with `as_dict()` and `[a.kind.value for a in parse(s).ambiguities]` over each string above and its `.upper()` and `.lower()`. +- 2026-10-01 (Derek), #554 — A WORD OF BOTH THE PARTICLE AND THE SUFFIX VOCABULARY COUNTS IN THE RUN, AND STILL SPEAKS FOR NOTHING. After #544 the run test (`_segment.py`) counts `vd` and `mc` as suffix words, while S2's company (`_pieces._anchors`) refuses to let them speak for the word behind them, an exclusion added within PR #552 after that PR's own anchor pass had split `Smith vd Ma, John` into family 'Smith Ma', suffix 'vd' (an intermediate state, never released). The issue framed that as two answers to one question and offered (b): end the run at such a particle, so that one "has a live name reading" predicate decides both. DECIDED (a), the shipped behavior, and no parse moves. They are two questions. Run membership asks what the word IS, and `vd` is unambiguous suffix vocabulary; company asks whether the word SPEAKS FOR the one behind it, and a particle cannot, P2 joining it forward to exactly that word. Both answers follow from the word's vocabulary, and they differ by design rather than by drift; S2's "end the run" is the company's run, not the part C1 counts. The clause reaches exactly `vd` and `mc`: `do`, also in both vocabularies, is in the ambiguous half and so is already a word of C1's class, reporting where these two do not (`John Smith, Jr do` against `John Smith, Jr vd`). Recompute the set: `L = Parser().lexicon; L.particles & (L.suffix_acronyms | L.suffix_words) - L.suffix_acronyms_ambiguous`, and drop the subtraction to see `do` with them. + The evidence is the writing that carries no other signal. Neither reading of `vd Ma` after a comma is realistic, a particle in front of a Chinese family name being as rare as `vd` or `mc` as a credential, so the mixed-case writing is an arbitrary edge case and decides nothing. Written in one case, where S2 has no capitals to read and the vocabulary is all there is, the part reads as the credential run, and that is the reading a reader would take: measured 2026-10-01 on master (fc682e36), `JOHN SMITH, VD MA` and `john smith, vd ma` read given 'JOHN'/'john', family 'SMITH'/'smith', suffix 'VD MA'/'vd ma' and report `suffix-or-name`, as `John Smith, vd Ma` does, and `John Smith, mc Ma` reads the same in every casing. It is also 1.4.0's reading, verified against the released wheel the same day: 1.4.0 read suffix 'vd Ma' and 'VD MA', and 2.0.0 through 2.3.0 read the whole name as the family. (b) would add a particle exception to C1, and to #563's paired-initials test besides, in order to give `JOHN SMITH, VD MA` family 'JOHN SMITH VD MA' with no given name: the worse reading, on exactly the input where nothing else speaks. + #563's site needs nothing of its own. `García Márquez, vd G.J.` reads as `García Márquez, PhD G.J.` does, in every casing: given 'García', family 'Márquez', suffix 'vd G.J.', and neither reports. If a surname split by a credential run in silence is a defect, it is the run's, not the particle's. Open: #573 — uniform-case `vd`/`mc` outside the run after a suffix comma. Recompute: `parse(s)` with `as_dict()` and `[a.kind.value for a in parse(s).ambiguities]` over each string above and its `.upper()` and `.lower()`. ### T1 — separators, not joiners diff --git a/docs/design/rules.md b/docs/design/rules.md index d400e464..a5680a26 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1703,15 +1703,17 @@ C1. Rationale: a credential run after the comma means the name is in case: one letter is the shape a middle initial is written in, and S3 retires single-character matches for the same reason, while a multi-letter numeral or generational word stands in a - run like any suffix word ('John Smith, III Ma'). A word of both - the particle and the unambiguous suffix vocabulary counts as a - suffix word there too ('John Smith, vd Ma'): membership in this - part asks what the word is, while S2's company asks whether it - speaks for the word behind it, which a particle never does, so - the company run it ends is S2's and not this one. A word of both - the title and the suffix vocabulary opening such a part counts as - a suffix word there, the name before the comma being complete, - except as the titles of paired initials, above. Where every word + run like any suffix word ('John Smith, III Ma'). A word's + particle reading does not take it out of the run: a word of both + the particle and the unambiguous suffix vocabulary is a suffix + word there too ('John Smith, vd Ma'). Run membership asks what + the word is, and S2's company asks whether it speaks for the word + behind it, which a particle never does; the run S2 says such a + word ends is the company's, not this count's. A word of both the + title and the suffix vocabulary opening a part the count reads as + a run counts as a suffix word there, the name before the comma + being complete, except as the titles of paired initials, above. + Where every word of this class in the part is a LISTED word written in capitals in a mixed-case name, the writing has already made each of them the credential (S2), and the part reads as the credential run on that @@ -1819,6 +1821,9 @@ C1. Rationale: a credential run after the comma means the name is in "John Smith, PhD vd DO" → suffix="PhD vd DO" "John Smith, vd Ma" → suffix="vd Ma" "JOHN SMITH, VD MA" → suffix="VD MA" + "JOHN SMITH, VD MA" → ambiguities=("suffix-or-name",) + "John Smith, Jr vd" → ambiguities=() + "John Smith, Jr do" → ambiguities=("suffix-or-name",) · boundary "John Smith, X.Y.Z." → suffix="X.Y.Z." "John Smith, X.Y.Z." → ambiguities=() "John Smith, X.Y.Z." unlisted_dotted_suffixes-off → given="X.Y.Z." diff --git a/tests/v2/test_ledger_guards.py b/tests/v2/test_ledger_guards.py index 93b66389..94b647e9 100644 --- a/tests/v2/test_ledger_guards.py +++ b/tests/v2/test_ledger_guards.py @@ -3071,7 +3071,8 @@ class _LatinCopy(NamedTuple): # suffix examples, a vocabulary word in front of a member, which # copies no set either. frozenset({"JOHN SMITH, VD MA", "John Smith, Ed Ma", - "John Smith, Ms Ma", "John Smith, PhD Ma", + "John Smith, Jr do", "John Smith, Ms Ma", + "John Smith, PhD Ma", r"John Smith, X\.Y\.Z\. MA", "John Smith, vd Ma", "john smith, md ma"}), frozenset({"Doe, Jane PhD MEng", "Doe, Jane nee Smith PhD MEng", @@ -3655,10 +3656,15 @@ def _claim(rule: dict) -> _Claim: # 2026-10-01, #562: 27 -> 29, 'John Smith, PhD DO DO' # and 'John Smith, PhD vd DO', #562's rules.md#C1 examples # and case rows. Reach, verified name by name. + # 2026-10-01, #554: 29 -> 31, 'John Smith, Jr vd' and 'John + # Smith, Jr do', #554's rules.md#C1 examples. Reach, verified + # name by name. "fix(#379) a tussenvoegsel after a family comma attaches to the family": - _Claim(29, ('family', 'middle'), "36879da9fc51", None), + _Claim(31, ('family', 'middle'), "233aa8786b15", None), + # 2026-10-01, #554: 2 -> 3, 'John Smith, Jr vd', #554's + # rules.md#C1 example. Reach, verified by name. "fix(#380) a trailing vd after a family comma is the tussenvoegsel, not a post-nominal": - _Claim(2, ('family', 'suffix'), "ec0d45289dc1", None), + _Claim(3, ('family', 'suffix'), "d5cd77f5b24b", None), # 279 -> 280 with #371, and the growth is corpus, not behavior: # that PR added `Ph. D., John` as a rules.md example, so the # regex matches one more corpus name. The name does not diff at @@ -3765,11 +3771,12 @@ def _claim(rule: dict) -> _Claim: # 2026-10-01, #562: 408 -> 410, 'John Smith, PhD DO DO' # and 'John Smith, PhD vd DO', #562's rules.md#C1 examples # and case rows. Reach, verified name by name. - # 2026-10-01, #554: 410 -> 412, 'John Smith, vd Ma' and - # 'JOHN SMITH, VD MA', #554's rules.md#C1 examples. Reach, - # verified name by name. + # 2026-10-01, #554: 410 -> 414, 'John Smith, vd Ma', + # 'JOHN SMITH, VD MA', 'John Smith, Jr vd' and 'John Smith, + # Jr do', #554's rules.md#C1 examples. Reach, verified name by + # name. "fix(comma-family) lone post-comma piece routes to suffix/title, not first": - _Claim(412, ('given', 'suffix', 'title'), "d9217453f0df", None), + _Claim(414, ('given', 'suffix', 'title'), "46663c4aa90a", None), "fix(comma-family) a comma followed only by titles keeps the given/family split": _Claim(2, ('family', 'given'), "5bd9c6d96c38", None), "fix(comma-family) a comma followed only by titles keeps the given/family split, the C1 example": @@ -3864,11 +3871,12 @@ def _claim(rule: dict) -> _Claim: # 2026-10-01, #562: 408 -> 410, 'John Smith, PhD DO DO' # and 'John Smith, PhD vd DO', #562's rules.md#C1 examples # and case rows. Reach, verified name by name. - # 2026-10-01, #554: 410 -> 412, 'John Smith, vd Ma' and - # 'JOHN SMITH, VD MA', #554's rules.md#C1 examples. Reach, - # verified name by name. + # 2026-10-01, #554: 410 -> 414, 'John Smith, vd Ma', + # 'JOHN SMITH, VD MA', 'John Smith, Jr vd' and 'John Smith, + # Jr do', #554's rules.md#C1 examples. Reach, verified name by + # name. "fix(comma-precomma-family) pre-comma run reads as family, not given": - _Claim(412, ('family', 'given'), "d9217453f0df", None), + _Claim(414, ('family', 'given'), "46663c4aa90a", None), # 2026-09-20, #397: retitled in place, reach and digest # unchanged -- the rule keeps 'Carod i', which the landing # leaves byte-identical. @@ -4658,7 +4666,10 @@ def _claim(rule: dict) -> _Claim: # 2026-10-01, #562: 27 -> 29, 'John Smith, PhD DO DO' # and 'John Smith, PhD vd DO', #562's rules.md#C1 examples # and case rows. Reach, verified name by name. - _Claim(29, ('_ambiguities', 'family', 'middle'), "36879da9fc51", None), + # 2026-10-01, #554: 29 -> 31, 'John Smith, Jr vd' and 'John + # Smith, Jr do', #554's rules.md#C1 examples. Reach, verified + # name by name. + _Claim(31, ('_ambiguities', 'family', 'middle'), "233aa8786b15", None), # 2026-09-18: 126 -> 131. Five corpus names arrived with # #289/#516's own case rows -- 'J.씨', 'John Smith 田.中.', # '毛泽东, MA', '田中 太郎, MA', '마틴 킹, MA' -- all of them @@ -4693,8 +4704,10 @@ def _claim(rule: dict) -> _Claim: _Claim(1, ('_ambiguities', 'family', 'given'), "ca7b37af6cf8", None), "fix(#367) a title no longer displaces a leading particle out of the leading position": _Claim(3, ('family', 'given'), "724967a4a117", None), + # 2026-10-01, #554: 2 -> 3, 'John Smith, Jr vd', #554's + # rules.md#C1 example. Reach, verified by name. "fix(#380) a trailing vd after a family comma is the tussenvoegsel, not a post-nominal": - _Claim(2, ('_ambiguities', 'family', 'suffix'), "ec0d45289dc1", None), + _Claim(3, ('_ambiguities', 'family', 'suffix'), "d5cd77f5b24b", None), "fix(#399) a maiden marker bounds the particle chain that swallowed it": _Claim(6, ('family', 'maiden'), "89e1f4afdd2a", ('DEFAULT',)), "fix(#360) mc moved into the never-given particles, so it folds into the family": @@ -5200,10 +5213,10 @@ def _claim(rule: dict) -> _Claim: # 2026-09-27, #544: new, 4; gains 'John Smith, Ed Ma', 'John # Smith, Ms Ma', 'John Smith, X.Y.Z. MA', 'john smith, md ma'. # 2026-09-28, #544: 4 -> 5; gains 'John Smith, PhD Ma'. - # 2026-10-01, #554: 5 -> 7; gains 'John Smith, vd Ma' and - # 'JOHN SMITH, VD MA'. + # 2026-10-01, #554: 5 -> 8; gains 'John Smith, vd Ma', + # 'JOHN SMITH, VD MA' and 'John Smith, Jr do'. "fix(#544) a run of ambiguous members after a suffix comma reads by the name-word count": - _Claim(7, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "890e7723c601", ('DEFAULT',)), + _Claim(8, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "ac1125df8691", ('DEFAULT',)), # 2026-09-27, #544: new, 4; gains 'Doe, Jane PhD MEng', 'Doe, # Jane nee Smith PhD MEng', 'Jane Doe nee Smith PhD MEng', # 'John Smith PhD MEng'. Relabelled the same day @@ -5652,10 +5665,10 @@ def _claim(rule: dict) -> _Claim: # 2026-09-27, #544: new, 4; gains 'John Smith, Ed Ma', 'John # Smith, Ms Ma', 'John Smith, X.Y.Z. MA', 'john smith, md ma'. # 2026-09-28, #544: 4 -> 5; gains 'John Smith, PhD Ma'. - # 2026-10-01, #554: 5 -> 7; gains 'John Smith, vd Ma' and - # 'JOHN SMITH, VD MA'. + # 2026-10-01, #554: 5 -> 8; gains 'John Smith, vd Ma', + # 'JOHN SMITH, VD MA' and 'John Smith, Jr do'. "fix(#544) a run of ambiguous members after a suffix comma reads by the name-word count": - _Claim(7, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "890e7723c601", ('DEFAULT',)), + _Claim(8, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "ac1125df8691", ('DEFAULT',)), # 2026-09-27, #544: new, 4; gains 'Doe, Jane PhD MEng', 'Doe, # Jane nee Smith PhD MEng', 'Jane Doe nee Smith PhD MEng', # 'John Smith PhD MEng'. Relabelled the same day @@ -5837,13 +5850,18 @@ def _claim(rule: dict) -> _Claim: # 2026-10-01, #562: 27 -> 29, 'John Smith, PhD DO DO' # and 'John Smith, PhD vd DO', #562's rules.md#C1 examples # and case rows. Reach, verified name by name. - _Claim(29, ('_ambiguities', 'family', 'middle'), "36879da9fc51", None), + # 2026-10-01, #554: 29 -> 31, 'John Smith, Jr vd' and 'John + # Smith, Jr do', #554's rules.md#C1 examples. Reach, verified + # name by name. + _Claim(31, ('_ambiguities', 'family', 'middle'), "233aa8786b15", None), "fix(#424) an unlisted abbreviation is as transparent as a listed title to the leading particle": _Claim(1, ('_ambiguities', 'family', 'given'), "ca7b37af6cf8", None), "fix(#367) a title no longer displaces a leading particle out of the leading position": _Claim(3, ('family', 'given'), "724967a4a117", None), + # 2026-10-01, #554: 2 -> 3, 'John Smith, Jr vd', #554's + # rules.md#C1 example. Reach, verified by name. "fix(#380) a trailing vd after a family comma is the tussenvoegsel, not a post-nominal": - _Claim(2, ('_ambiguities', 'family', 'suffix'), "ec0d45289dc1", None), + _Claim(3, ('_ambiguities', 'family', 'suffix'), "d5cd77f5b24b", None), "fix(#399) a maiden marker bounds the particle chain that swallowed it": _Claim(6, ('family', 'maiden'), "89e1f4afdd2a", ('DEFAULT',)), "fix(#360) mc moved into the never-given particles, so it folds into the family": @@ -6338,10 +6356,10 @@ def _claim(rule: dict) -> _Claim: # 2026-09-27, #544: new, 4; gains 'John Smith, Ed Ma', 'John # Smith, Ms Ma', 'John Smith, X.Y.Z. MA', 'john smith, md ma'. # 2026-09-28, #544: 4 -> 5; gains 'John Smith, PhD Ma'. - # 2026-10-01, #554: 5 -> 7; gains 'John Smith, vd Ma' and - # 'JOHN SMITH, VD MA'. + # 2026-10-01, #554: 5 -> 8; gains 'John Smith, vd Ma', + # 'JOHN SMITH, VD MA' and 'John Smith, Jr do'. "fix(#544) a run of ambiguous members after a suffix comma reads by the name-word count": - _Claim(7, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "890e7723c601", ('DEFAULT',)), + _Claim(8, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "ac1125df8691", ('DEFAULT',)), # 2026-09-27, #544: new, 4; gains 'Doe, Jane PhD MEng', 'Doe, # Jane nee Smith PhD MEng', 'Jane Doe nee Smith PhD MEng', # 'John Smith PhD MEng'. Relabelled the same day @@ -6639,10 +6657,10 @@ def _claim(rule: dict) -> _Claim: # 2026-09-27, #544: new, 4; gains 'John Smith, Ed Ma', 'John # Smith, Ms Ma', 'John Smith, X.Y.Z. MA', 'john smith, md ma'. # 2026-09-28, #544: 4 -> 5; gains 'John Smith, PhD Ma'. - # 2026-10-01, #554: 5 -> 7; gains 'John Smith, vd Ma' and - # 'JOHN SMITH, VD MA'. + # 2026-10-01, #554: 5 -> 8; gains 'John Smith, vd Ma', + # 'JOHN SMITH, VD MA' and 'John Smith, Jr do'. "fix(#544) a run of ambiguous members after a suffix comma reads by the name-word count": - _Claim(7, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "890e7723c601", ('DEFAULT',)), + _Claim(8, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "ac1125df8691", ('DEFAULT',)), # 2026-09-27, #544: new, 4; gains 'Doe, Jane PhD MEng', 'Doe, # Jane nee Smith PhD MEng', 'Jane Doe nee Smith PhD MEng', # 'John Smith PhD MEng'. diff --git a/tools/differential/corpus_rules.jsonl b/tools/differential/corpus_rules.jsonl index 0d266b23..d816b00b 100644 --- a/tools/differential/corpus_rules.jsonl +++ b/tools/differential/corpus_rules.jsonl @@ -206,6 +206,8 @@ "John Smith, Ed" "John Smith, Ed Ma" "John Smith, Jones" +"John Smith, Jr do" +"John Smith, Jr vd" "John Smith, LEED AP" "John Smith, MA" "John Smith, MD, Bart" diff --git a/tools/differential/expected_since_2.0.0.toml b/tools/differential/expected_since_2.0.0.toml index 61a39228..d7268994 100644 --- a/tools/differential/expected_since_2.0.0.toml +++ b/tools/differential/expected_since_2.0.0.toml @@ -3769,15 +3769,19 @@ issue = "fix(#544) a run of ambiguous members after a suffix comma reads by the # like the rest. # # 2026-10-01, #554: 'John Smith, vd Ma' and 'JOHN SMITH, VD MA' join, -# rules.md#C1's examples for a word of both the particle and the suffix -# vocabulary, which "counts as a suffix word there too". This baseline -# read the whole name as the family ('John Smith vd Ma'). +# rules.md#C1's examples for a word of both the particle and the +# unambiguous suffix vocabulary: "A word's particle reading does not +# take it out of the run". This baseline read the whole name as the +# family ('John Smith vd Ma'). 'John Smith, Jr do' joins as C1's +# boundary for that clause: `do` is a word of the ambiguous class, so +# the count flips the part and reports, where this baseline read +# title 'Jr', given 'do'. # # Literal; `fields` is the union the names move ('title' is the duals'). # Probes: 'Smith, Ed Ma' (one name word before the comma keeps the # listing form), 'Smith, Ms Ma' (a dual opening the given part is a # title) and the superstring 'Dr. John Smith, Ed Ma' are _MUST_NOT_MATCH. -name_regex = "^(?:JOHN SMITH, VD MA|John Smith, Ed Ma|John Smith, Ms Ma|John Smith, PhD Ma|John Smith, X\\.Y\\.Z\\. MA|John Smith, vd Ma|john smith, md ma)$" +name_regex = "^(?:JOHN SMITH, VD MA|John Smith, Ed Ma|John Smith, Jr do|John Smith, Ms Ma|John Smith, PhD Ma|John Smith, X\\.Y\\.Z\\. MA|John Smith, vd Ma|john smith, md ma)$" fields = ["given", "middle", "family", "suffix", "title", "_ambiguities"] orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.1.0.toml b/tools/differential/expected_since_2.1.0.toml index 1e624fe9..291a4cf7 100644 --- a/tools/differential/expected_since_2.1.0.toml +++ b/tools/differential/expected_since_2.1.0.toml @@ -3680,15 +3680,19 @@ issue = "fix(#544) a run of ambiguous members after a suffix comma reads by the # like the rest. # # 2026-10-01, #554: 'John Smith, vd Ma' and 'JOHN SMITH, VD MA' join, -# rules.md#C1's examples for a word of both the particle and the suffix -# vocabulary, which "counts as a suffix word there too". This baseline -# read the whole name as the family ('John Smith vd Ma'). +# rules.md#C1's examples for a word of both the particle and the +# unambiguous suffix vocabulary: "A word's particle reading does not +# take it out of the run". This baseline read the whole name as the +# family ('John Smith vd Ma'). 'John Smith, Jr do' joins as C1's +# boundary for that clause: `do` is a word of the ambiguous class, so +# the count flips the part and reports, where this baseline read +# title 'Jr', given 'do'. # # Literal; `fields` is the union the names move ('title' is the duals'). # Probes: 'Smith, Ed Ma' (one name word before the comma keeps the # listing form), 'Smith, Ms Ma' (a dual opening the given part is a # title) and the superstring 'Dr. John Smith, Ed Ma' are _MUST_NOT_MATCH. -name_regex = "^(?:JOHN SMITH, VD MA|John Smith, Ed Ma|John Smith, Ms Ma|John Smith, PhD Ma|John Smith, X\\.Y\\.Z\\. MA|John Smith, vd Ma|john smith, md ma)$" +name_regex = "^(?:JOHN SMITH, VD MA|John Smith, Ed Ma|John Smith, Jr do|John Smith, Ms Ma|John Smith, PhD Ma|John Smith, X\\.Y\\.Z\\. MA|John Smith, vd Ma|john smith, md ma)$" fields = ["given", "middle", "family", "suffix", "title", "_ambiguities"] orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.2.0.toml b/tools/differential/expected_since_2.2.0.toml index 784ff85a..bc2b8008 100644 --- a/tools/differential/expected_since_2.2.0.toml +++ b/tools/differential/expected_since_2.2.0.toml @@ -2078,15 +2078,19 @@ issue = "fix(#544) a run of ambiguous members after a suffix comma reads by the # like the rest. # # 2026-10-01, #554: 'John Smith, vd Ma' and 'JOHN SMITH, VD MA' join, -# rules.md#C1's examples for a word of both the particle and the suffix -# vocabulary, which "counts as a suffix word there too". This baseline -# read the whole name as the family ('John Smith vd Ma'). +# rules.md#C1's examples for a word of both the particle and the +# unambiguous suffix vocabulary: "A word's particle reading does not +# take it out of the run". This baseline read the whole name as the +# family ('John Smith vd Ma'). 'John Smith, Jr do' joins as C1's +# boundary for that clause: `do` is a word of the ambiguous class, so +# the count flips the part and reports, where this baseline read +# given 'Jr', family 'do John Smith'. # # Literal; `fields` is the union the names move ('title' is the duals'). # Probes: 'Smith, Ed Ma' (one name word before the comma keeps the # listing form), 'Smith, Ms Ma' (a dual opening the given part is a # title) and the superstring 'Dr. John Smith, Ed Ma' are _MUST_NOT_MATCH. -name_regex = "^(?:JOHN SMITH, VD MA|John Smith, Ed Ma|John Smith, Ms Ma|John Smith, PhD Ma|John Smith, X\\.Y\\.Z\\. MA|John Smith, vd Ma|john smith, md ma)$" +name_regex = "^(?:JOHN SMITH, VD MA|John Smith, Ed Ma|John Smith, Jr do|John Smith, Ms Ma|John Smith, PhD Ma|John Smith, X\\.Y\\.Z\\. MA|John Smith, vd Ma|john smith, md ma)$" fields = ["given", "middle", "family", "suffix", "title", "_ambiguities"] orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.3.0.toml b/tools/differential/expected_since_2.3.0.toml index 99ee3e58..850ac259 100644 --- a/tools/differential/expected_since_2.3.0.toml +++ b/tools/differential/expected_since_2.3.0.toml @@ -1345,15 +1345,19 @@ issue = "fix(#544) a run of ambiguous members after a suffix comma reads by the # like the rest. # # 2026-10-01, #554: 'John Smith, vd Ma' and 'JOHN SMITH, VD MA' join, -# rules.md#C1's examples for a word of both the particle and the suffix -# vocabulary, which "counts as a suffix word there too". This baseline -# read the whole name as the family ('John Smith vd Ma'). +# rules.md#C1's examples for a word of both the particle and the +# unambiguous suffix vocabulary: "A word's particle reading does not +# take it out of the run". This baseline read the whole name as the +# family ('John Smith vd Ma'). 'John Smith, Jr do' joins as C1's +# boundary for that clause: `do` is a word of the ambiguous class, so +# the count flips the part and reports, where this baseline read +# given 'Jr', family 'do John Smith'. # # Literal; `fields` is the union the names move ('title' is the duals'). # Probes: 'Smith, Ed Ma' (one name word before the comma keeps the # listing form), 'Smith, Ms Ma' (a dual opening the given part is a # title) and the superstring 'Dr. John Smith, Ed Ma' are _MUST_NOT_MATCH. -name_regex = "^(?:JOHN SMITH, VD MA|John Smith, Ed Ma|John Smith, Ms Ma|John Smith, PhD Ma|John Smith, X\\.Y\\.Z\\. MA|John Smith, vd Ma|john smith, md ma)$" +name_regex = "^(?:JOHN SMITH, VD MA|John Smith, Ed Ma|John Smith, Jr do|John Smith, Ms Ma|John Smith, PhD Ma|John Smith, X\\.Y\\.Z\\. MA|John Smith, vd Ma|john smith, md ma)$" fields = ["given", "middle", "family", "suffix", "title", "_ambiguities"] orders = ["DEFAULT"] From 372d5ec1e62596a19ec7a1f1d154319dc3792cb3 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Thu, 1 Oct 2026 18:16:33 -0700 Subject: [PATCH 4/4] docs(C1): second review round for #554 -- history, Open handle, labels MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - decisions.md#C1: the 'Smith vd Ma, John' split was not only an intermediate state of #552 -- 2.2.0 and 2.3.0 shipped it and #530 removed it this cycle (measured against the wheels). - The Open: handle names the shapes #573 actually covers: before a comma and after a family comma, not after a suffix comma. - 'John Smith, Jr do' is not a boundary of the particle clause: do is still a run member and the part still flips; only the report differs. The `· boundary` label and the ledgers' "boundary" wording go. Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 4 ++-- docs/design/rules.md | 2 +- tools/differential/expected_since_2.0.0.toml | 8 ++++---- tools/differential/expected_since_2.1.0.toml | 8 ++++---- tools/differential/expected_since_2.2.0.toml | 8 ++++---- tools/differential/expected_since_2.3.0.toml | 8 ++++---- 6 files changed, 19 insertions(+), 19 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 62001bc4..8418c3b8 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -877,9 +877,9 @@ Excluded (MAIDEN_MARKERS, per nameparser/config/maiden_markers.py): - 2026-09-28 (Derek), #544 — A PART READ WHOLLY AS SUFFIXES REPORTS NO NAME READING OF ITS WORDS. group's particle chain runs over every comma segment, and its two emitters — `particle-or-given` when a particle behind a word of both the title and the particle vocabulary chains (since 2.0.0, de264af1) and `suffix-or-name` when the chain takes an ambiguous acronym into the name (#289/#516, 59d8f38a, in no release) — reported inside a TAIL segment, which assign reads wholly as suffixes. Each such report named a reading the parse never made, against rules.md#A1's "A report names the reading the parse took". A family comma's tail was already silent, since group hands the chain no report list anywhere after a family comma; the suffix comma's tails were not. Measured on the released wheels: `John Smith, Jr., Freiherr von Richthofen` reports `particle-or-given` on 'von', a token in the suffix role, at 2.0.0, 2.1.0, 2.2.0 and 2.3.0; `John Smith, Jr., PhD van Ma` and `John Smith, Jr., PhD Do Ma` report it on 'van' and 'Do' at 2.0.0 and 2.1.0 only. The `suffix-or-name` half reached `John Smith, Jr., PhD Do Ma`, `John Smith, MA, PhD Do Ma` and `John Smith, Jr., PhD van Ma` on master (e10e83b4), and the run rule of the bullet above made it reachable behind ONE comma: `John Smith, PhD Do Ma` carried C1's flip and a second report on 'Ma'. FIXED by scope, not by a new test at the emitter: group passes the chain no report list in a tail segment either, so both emitters go quiet there together, and rules.md#C2 states the boundary for any part consumed wholly as suffixes. The maiden channel is a separate parameter and is untouched: a tail segment's reader is NONE, so the maiden walk reports nothing there to begin with. No mechanisms.md entry: this is AMBIGUITY-AT-THE-DECISION-SITE's own contract (a report fires only where the parse chose between live readings) applied to a stage whose reading a later stage overrides for the whole segment. MEASURED 2026-09-28, the tree against the same tree with `None if family_comma else ambiguities` restored in `group()` (the comparator), each parse recorded as its seven fields plus `(kind, [(token text, token role)])` per report, under all three name orders: 0 of the 1441 differential-corpus names move, so the gate has nothing to classify; over a comma grid — the prefixes `John Smith, `, `John Smith, Jr., ` and `Smith, John, ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, Do, van, de, Jr, MEng, Ed, y, i}, each text as written, lowercased and uppercased, deduplicated to 11,049 texts — 360 parses (120 texts, every one of them mixed case) lose one `suffix-or-name` report apiece, every removed report on a token in the suffix role, 0 reports added, 0 field moves. 21 of the 120 texts carry one comma and are the run rule's reach; the other 99 carry two and moved the same way (297 parses) when the same one-line change was applied to master e10e83b4; the `Smith, John, ` prefix moves nothing, being a family comma. The grid holds no word of both the title and the particle vocabulary, so the `particle-or-given` half is witnessed by the case row alone. Pinned by the case rows `a_credential_run_after_the_comma_reports_no_chain_fork` and `a_part_past_the_second_reports_no_particle_fork`; the rows the chain still reports on outside a tail are `the_chain_reports_the_acronym_it_takes` and `titled_particle_chain_survives_a_title_that_is_also_a_particle`. - 2026-10-01 (Derek), #562 — A PARTICLE CHAIN UNSETTLES A RUN THE CAPITALS SETTLED, AND THE COUNT READS IT. The 2026-09-27 bullet above left a run whose every member is listed and leans credential to the family-comma path, "which already reads it whole". That promise fails wherever two particles stand side by side in the part: group chains them into one particle run (P2), assign reads the part as name text, and P6 attaches the chain to the family — `John Smith, PhD DO DO` read given 'PhD', family 'DO DO John Smith', and `John Smith, DO DO DO` given 'DO', middle 'DO DO'. rules.md#S2 already said the capitals do not decide a member chained behind another particle ("the run attaches whatever the capitals say (P6)"), so C1's shortcut was resting on a premise S2 denies. Of the two fixes #562 weighed, the one taken narrows the shortcut and leaves S2 as written: a part holding two particles side by side is read by the count, which two name words before the comma flip to the credential run, reported (`suffix-or-name`). The other — letting C1's evidence or S2's credential-in-front company outrank the chain — would have contradicted S2's sentence and P6's `Doe, John van DO` example, so it needed S2 amended rather than a gap filled. The test is ANY two adjacent particles, not a member behind one: `vd` is a particle and an unambiguous suffix word, so `John Smith, PhD vd DO` and `John Smith, MA vd vd` chained and misread the same way, the second with no member behind a particle at all. Segment runs before classify, so it asks classify's own predicate (`_normalize(text) in lexicon.particles`) and only while the run is still settled. The test does not ask whether the family-comma path would actually have misread the part, which it could not without reading ahead to group: where that path did read the part whole — for example a pair opening the part with an unambiguous particle-and-suffix word (`John Smith, vd DO`, `John Smith, VD DO`), a credential that is also a title in front (`John Smith, MD DO DO`), or a credential closed by a period in front (`John Smith, Esq. DO DO`, `John Smith, Jr. DO DO`), the list being by example rather than a census — the count flips the part to the same fields and reports the call, as every flip at this comma does (rules.md#C1's "A decision either way at this comma is reported"). Those reports are ACCEPTED (Derek, 2026-10-01: none of these is a name anyone would write on purpose, so a report is the right signal): with the capitals no longer settling the run, the call is the count's, and the report says so; `tests/v2/cases.py` pins `John Smith, vd DO`. 1.4.0 read every one of these names as the fix does; 2.0.0 and 2.1.0 read `John Smith, PhD DO DO` as title 'PhD', given 'DO DO', and 2.2.0 and 2.3.0 as the issue describes. MEASURED 2026-10-01 against master 0eadedeb, py3.11, `nameparser.__file__` asserted on each side, each parse compared as its seven fields plus its sorted ambiguity kinds: 0 of the 1453 differential-corpus names move (the two names this change adds are the gate's only movers, under a new fix(#562) rule in the four 2.x ledgers); over tests/v2/test_properties.py's settled grid (5,580 texts) 15 move, 10 of them role moves, and the other 5 (`John Smith, MD DO DO`, `MS`, `Esq.`, `Sr`, `Ms` in front) keep their fields and gain the flip's report; over a wider grid — the prefixes `John Smith, `, `Smith, ` and `Doe, John ` times every run of one to three words drawn with repetition from {PhD, MA, Ma, DO, Do, do, vd, van, Jr, MD, Ms}, 4,389 texts — 30 move, 17 of them role moves, every mover a `John Smith, ` text now reading given 'John', family 'Smith' and the whole part as suffix, and every one reporting `suffix-or-name`. Recompute: check out the parent into a separate worktree, parse each grid in both trees under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff. The settled grid's own pin moves with it: tests/v2/test_properties.py's `_SETTLED_COUNT` reads 2,994 where it read 3,024, the ten `_SETTLED_EXCEPTIONS` that pinned #562 are gone, and its two recorded negative controls read 1,142 and 215 — the second had already moved from 751 to 215 with #563, before this change. LEFT OPEN, as #562 asked: `Smith, PhD DO DO` (ONE name word before the comma, so the count keeps the listing form, and P6 attaches the chain: given 'PhD', family 'DO DO Smith'), where no rule states S2's credential-in-front company against a chain; and `John Smith, PhD van der`, whose particles are no members and never reach the run test, reading given 'PhD', family 'van der John Smith' as before. - 2026-10-01 (Derek) — A CAPITALS-SETTLED RUN OPENED BY A WORD OF THE CLASS REPORTS, AND RULES.MD#C1 NOW SAYS SO. C1 said a run whose every class word is written in capitals "reads whole in silence". That held only behind another credential: a class word OPENING the part is the first word after the comma, whose decision C1 reports either way, so `Doe, MA PhD`, `John Smith, MA MA` and `Smith, MA PhD` read wholly as suffixes and report `suffix-or-name` once, on that word, while `John Smith, PhD MA` and `Smith, PhD MA` are silent. The behavior predates #562 and is kept: nobody repeats `MA` at the end of their name on purpose, so the report is the right signal. Statement corrected, no parse moved. MEASURED 2026-10-01 on master b39c370c, the prefixes `John Smith, ` and `Smith, ` times every run of two or three words drawn with repetition from {MA, BA, ED, DO, JD, PhD, MD, Jr, Esq.} holding at least one class word: all 900 runs opening with a class word report, 895 of them read whole (14 of those hold a #562 particle pair, whose count flip reports over the whole part rather than on the first word) and the other five being `Smith, MA DO DO` and its like, the one-name-word chain #562 left open; of the 560 opening with another credential, 554 read whole in silence, and the other six hold a #562 particle pair — four reading whole and reporting (`John Smith, PhD DO DO`), two being `Smith, PhD DO DO` and `Smith, Jr DO DO`. Recompute: parse that grid and bucket by whether the first word is a class word, whether `ambiguities` is empty, and whether the suffix is the whole part. The same day (Derek) rules.md#S2's first-slot precedence sentence gained the pointer to that exception, so the next reader does not rediscover it: S2 says the count decides the first slot after a family comma before case, true for one word (`John Smith, MA` flips by the count) and false for a capitals-settled run, where C1's shortcut reads the capitals first and the comma keeps its family reading. And S2's list of reporting slots lost "the segments beyond it": no `suffix-or-name` report lands past the second comma (`Smith, John, MA` and `John Smith, Jr., MA` are silent). It once did: at #530's merge cc78c960 `John Smith, Jr., PhD Do Ma` reported `suffix-or-name` on 'Ma' beside its `comma-structure` flag, from group's chain emitter, which the 2026-09-28 #544 bullet above silenced in a tail. S2 now points at C2 for what such a part does report, which is not only `comma-structure` — an unclosed delimiter there adds `unbalanced-delimiter` (`John Smith, Jr., (Bob`), and a maiden clause standing in it is read as one (`Smith, John, Jr nee Jones MA`, maiden 'Jones MA'). Derek, 2026-10-01: the silence is right. -- 2026-10-01 (Derek), #554 — A WORD OF BOTH THE PARTICLE AND THE SUFFIX VOCABULARY COUNTS IN THE RUN, AND STILL SPEAKS FOR NOTHING. After #544 the run test (`_segment.py`) counts `vd` and `mc` as suffix words, while S2's company (`_pieces._anchors`) refuses to let them speak for the word behind them, an exclusion added within PR #552 after that PR's own anchor pass had split `Smith vd Ma, John` into family 'Smith Ma', suffix 'vd' (an intermediate state, never released). The issue framed that as two answers to one question and offered (b): end the run at such a particle, so that one "has a live name reading" predicate decides both. DECIDED (a), the shipped behavior, and no parse moves. They are two questions. Run membership asks what the word IS, and `vd` is unambiguous suffix vocabulary; company asks whether the word SPEAKS FOR the one behind it, and a particle cannot, P2 joining it forward to exactly that word. Both answers follow from the word's vocabulary, and they differ by design rather than by drift; S2's "end the run" is the company's run, not the part C1 counts. The clause reaches exactly `vd` and `mc`: `do`, also in both vocabularies, is in the ambiguous half and so is already a word of C1's class, reporting where these two do not (`John Smith, Jr do` against `John Smith, Jr vd`). Recompute the set: `L = Parser().lexicon; L.particles & (L.suffix_acronyms | L.suffix_words) - L.suffix_acronyms_ambiguous`, and drop the subtraction to see `do` with them. +- 2026-10-01 (Derek), #554 — A WORD OF BOTH THE PARTICLE AND THE SUFFIX VOCABULARY COUNTS IN THE RUN, AND STILL SPEAKS FOR NOTHING. After #544 the run test (`_segment.py`) counts `vd` and `mc` as suffix words, while S2's company (`_pieces._anchors`) refuses to let them speak for the word behind them, an exclusion added within PR #552 after that PR's own anchor pass had brought back a split of `Smith vd Ma, John` into family 'Smith Ma', suffix 'vd'. That split is not only an intermediate state of #552: 2.2.0 and 2.3.0 shipped it (2.0.0 and 2.1.0 read family 'Smith vd Ma'), #530 removed it in this cycle, and #552 kept it removed. The issue framed that as two answers to one question and offered (b): end the run at such a particle, so that one "has a live name reading" predicate decides both. DECIDED (a), the shipped behavior, and no parse moves. They are two questions. Run membership asks what the word IS, and `vd` is unambiguous suffix vocabulary; company asks whether the word SPEAKS FOR the one behind it, and a particle cannot, P2 joining it forward to exactly that word. Both answers follow from the word's vocabulary, and they differ by design rather than by drift; S2's "end the run" is the company's run, not the part C1 counts. The clause reaches exactly `vd` and `mc`: `do`, also in both vocabularies, is in the ambiguous half and so is already a word of C1's class, reporting where these two do not (`John Smith, Jr do` against `John Smith, Jr vd`). Recompute the set: `L = Parser().lexicon; L.particles & (L.suffix_acronyms | L.suffix_words) - L.suffix_acronyms_ambiguous`, and drop the subtraction to see `do` with them. The evidence is the writing that carries no other signal. Neither reading of `vd Ma` after a comma is realistic, a particle in front of a Chinese family name being as rare as `vd` or `mc` as a credential, so the mixed-case writing is an arbitrary edge case and decides nothing. Written in one case, where S2 has no capitals to read and the vocabulary is all there is, the part reads as the credential run, and that is the reading a reader would take: measured 2026-10-01 on master (fc682e36), `JOHN SMITH, VD MA` and `john smith, vd ma` read given 'JOHN'/'john', family 'SMITH'/'smith', suffix 'VD MA'/'vd ma' and report `suffix-or-name`, as `John Smith, vd Ma` does, and `John Smith, mc Ma` reads the same in every casing. It is also 1.4.0's reading, verified against the released wheel the same day: 1.4.0 read suffix 'vd Ma' and 'VD MA', and 2.0.0 through 2.3.0 read the whole name as the family. (b) would add a particle exception to C1, and to #563's paired-initials test besides, in order to give `JOHN SMITH, VD MA` family 'JOHN SMITH VD MA' with no given name: the worse reading, on exactly the input where nothing else speaks. - #563's site needs nothing of its own. `García Márquez, vd G.J.` reads as `García Márquez, PhD G.J.` does, in every casing: given 'García', family 'Márquez', suffix 'vd G.J.', and neither reports. If a surname split by a credential run in silence is a defect, it is the run's, not the particle's. Open: #573 — uniform-case `vd`/`mc` outside the run after a suffix comma. Recompute: `parse(s)` with `as_dict()` and `[a.kind.value for a in parse(s).ambiguities]` over each string above and its `.upper()` and `.lower()`. + #563's site needs nothing of its own. `García Márquez, vd G.J.` reads as `García Márquez, PhD G.J.` does, in every casing: given 'García', family 'Márquez', suffix 'vd G.J.', and neither reports. If a surname split by a credential run in silence is a defect, it is the run's, not the particle's. Open: #573 — uniform-case `vd`/`mc` before a comma and after a family comma, and the silent mixed-case `Doe, Jane PhD vd Ma`. Recompute: `parse(s)` with `as_dict()` and `[a.kind.value for a in parse(s).ambiguities]` over each string above and its `.upper()` and `.lower()`. ### T1 — separators, not joiners diff --git a/docs/design/rules.md b/docs/design/rules.md index a5680a26..d55988ff 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1823,7 +1823,7 @@ C1. Rationale: a credential run after the comma means the name is in "JOHN SMITH, VD MA" → suffix="VD MA" "JOHN SMITH, VD MA" → ambiguities=("suffix-or-name",) "John Smith, Jr vd" → ambiguities=() - "John Smith, Jr do" → ambiguities=("suffix-or-name",) · boundary + "John Smith, Jr do" → ambiguities=("suffix-or-name",) "John Smith, X.Y.Z." → suffix="X.Y.Z." "John Smith, X.Y.Z." → ambiguities=() "John Smith, X.Y.Z." unlisted_dotted_suffixes-off → given="X.Y.Z." diff --git a/tools/differential/expected_since_2.0.0.toml b/tools/differential/expected_since_2.0.0.toml index d7268994..209325e5 100644 --- a/tools/differential/expected_since_2.0.0.toml +++ b/tools/differential/expected_since_2.0.0.toml @@ -3772,10 +3772,10 @@ issue = "fix(#544) a run of ambiguous members after a suffix comma reads by the # rules.md#C1's examples for a word of both the particle and the # unambiguous suffix vocabulary: "A word's particle reading does not # take it out of the run". This baseline read the whole name as the -# family ('John Smith vd Ma'). 'John Smith, Jr do' joins as C1's -# boundary for that clause: `do` is a word of the ambiguous class, so -# the count flips the part and reports, where this baseline read -# title 'Jr', given 'do'. +# family ('John Smith vd Ma'). 'John Smith, Jr do' joins as the +# contrast to the silent 'John Smith, Jr vd': `do` is a word of the +# ambiguous class, so the count flips the part and reports, where this +# baseline read title 'Jr', given 'do'. # # Literal; `fields` is the union the names move ('title' is the duals'). # Probes: 'Smith, Ed Ma' (one name word before the comma keeps the diff --git a/tools/differential/expected_since_2.1.0.toml b/tools/differential/expected_since_2.1.0.toml index 291a4cf7..d96f4faf 100644 --- a/tools/differential/expected_since_2.1.0.toml +++ b/tools/differential/expected_since_2.1.0.toml @@ -3683,10 +3683,10 @@ issue = "fix(#544) a run of ambiguous members after a suffix comma reads by the # rules.md#C1's examples for a word of both the particle and the # unambiguous suffix vocabulary: "A word's particle reading does not # take it out of the run". This baseline read the whole name as the -# family ('John Smith vd Ma'). 'John Smith, Jr do' joins as C1's -# boundary for that clause: `do` is a word of the ambiguous class, so -# the count flips the part and reports, where this baseline read -# title 'Jr', given 'do'. +# family ('John Smith vd Ma'). 'John Smith, Jr do' joins as the +# contrast to the silent 'John Smith, Jr vd': `do` is a word of the +# ambiguous class, so the count flips the part and reports, where this +# baseline read title 'Jr', given 'do'. # # Literal; `fields` is the union the names move ('title' is the duals'). # Probes: 'Smith, Ed Ma' (one name word before the comma keeps the diff --git a/tools/differential/expected_since_2.2.0.toml b/tools/differential/expected_since_2.2.0.toml index bc2b8008..c99de14a 100644 --- a/tools/differential/expected_since_2.2.0.toml +++ b/tools/differential/expected_since_2.2.0.toml @@ -2081,10 +2081,10 @@ issue = "fix(#544) a run of ambiguous members after a suffix comma reads by the # rules.md#C1's examples for a word of both the particle and the # unambiguous suffix vocabulary: "A word's particle reading does not # take it out of the run". This baseline read the whole name as the -# family ('John Smith vd Ma'). 'John Smith, Jr do' joins as C1's -# boundary for that clause: `do` is a word of the ambiguous class, so -# the count flips the part and reports, where this baseline read -# given 'Jr', family 'do John Smith'. +# family ('John Smith vd Ma'). 'John Smith, Jr do' joins as the +# contrast to the silent 'John Smith, Jr vd': `do` is a word of the +# ambiguous class, so the count flips the part and reports, where this +# baseline read given 'Jr', family 'do John Smith'. # # Literal; `fields` is the union the names move ('title' is the duals'). # Probes: 'Smith, Ed Ma' (one name word before the comma keeps the diff --git a/tools/differential/expected_since_2.3.0.toml b/tools/differential/expected_since_2.3.0.toml index 850ac259..c6103486 100644 --- a/tools/differential/expected_since_2.3.0.toml +++ b/tools/differential/expected_since_2.3.0.toml @@ -1348,10 +1348,10 @@ issue = "fix(#544) a run of ambiguous members after a suffix comma reads by the # rules.md#C1's examples for a word of both the particle and the # unambiguous suffix vocabulary: "A word's particle reading does not # take it out of the run". This baseline read the whole name as the -# family ('John Smith vd Ma'). 'John Smith, Jr do' joins as C1's -# boundary for that clause: `do` is a word of the ambiguous class, so -# the count flips the part and reports, where this baseline read -# given 'Jr', family 'do John Smith'. +# family ('John Smith vd Ma'). 'John Smith, Jr do' joins as the +# contrast to the silent 'John Smith, Jr vd': `do` is a word of the +# ambiguous class, so the count flips the part and reports, where this +# baseline read given 'Jr', family 'do John Smith'. # # Literal; `fields` is the union the names move ('title' is the duals'). # Probes: 'Smith, Ed Ma' (one name word before the comma keeps the