diff --git a/docs/customize.rst b/docs/customize.rst index 54430543..3a2d5aa3 100644 --- a/docs/customize.rst +++ b/docs/customize.rst @@ -502,7 +502,15 @@ listed below. ending a maiden marker's clause, also since 2.4: ``"Jane Doe nee Smith X.Y.Z."`` gives maiden ``Smith`` with suffix ``X.Y.Z.``, where ``False`` keeps maiden - ``Smith X.Y.Z.``. + ``Smith X.Y.Z.``. Two single letters right after a comma are + the exception: they are how a person's initials are written, + and the words before the comma may be one surname of two + words, so ``"García Márquez, G.J."`` keeps given ``G.J.`` + (reported) unless an unambiguous post-nominal in front of them + that is not also a title, or another unlisted dotted word beside + them, says otherwise (``"John Smith, PhD X.Y."`` gives suffix + ``PhD X.Y.``, while ``"García Márquez, Ms G.J."`` gives title + ``Ms``, given ``G.J.``). Case is irrelevant — the periods are the signal. Whole-token vocabulary still wins (``M.A.``, ``Ph.D.``), and a single trailing period is not this shape diff --git a/docs/design/decisions.md b/docs/design/decisions.md index fc81e48a..a4140b52 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -681,6 +681,11 @@ for n in ('Smith, John','Smith, XYZ'): print(n, calls_for(off.parse, n), calls_f ADDENDUM 2026-09-28 (Derek). (e) A PARTICLE ANCHORS NOTHING, a fifth boundary beside the four WHAT MAY ANCHOR lists: a word of both the particle and the unambiguous suffix vocabulary (`vd` and `mc` in the shipped lexicon; `do` sits in the ambiguous half, so it is a member and never an anchor) heads the family name behind it (P2), and anchoring there split the family around a suffix — `Smith vd Ma, John` read family 'Smith Ma', suffix 'vd'; `Jan vd Ma` suffix 'vd Ma' and no family; `Smith Mc Ma, John` and `D. Mc Ba Ed, Smith` the same way. Each reads as e10e83b4 reads it again (family 'Smith vd Ma', 'vd Ma', 'Smith Mc Ma', 'D. Mc Ba Ed'). The exclusion is of the word in front only: a particle MEMBER behind a credential is still spoken for (`doe, jane v phd do` keeps suffix 'v phd do'), and a C1 run holding the particle is C1's reading rather than the anchor's (`Jan Berg, vd Ma` reads suffix 'vd Ma', as `Jan Berg, vd` reads suffix 'vd'). Measured over the grid of MEASURED above with `vd` and `Mc` added to its words (127,578 texts, 382,734 parses), against the tree before this addendum: 9,426 parses (3,142 texts) move, every one holding `vd` or `Mc` directly in front of a member; 8,940 return to e10e83b4's reading exactly, and the other 486 all hold `M.Eng.`, whose own change is (3) — 406 of them take e10e83b4's fields and differ only in the pick `M.Eng.` no longer reports, and 80 take the fields e10e83b4 gives the same text with `PhD` in `M.Eng.`'s place (`Jane Doe nee Smith M.Eng. vd Ma` reads middle 'Doe M.Eng.', family 'vd Ma', as `Jane Doe nee Smith PhD vd Ma` reads middle 'Doe PhD'). The differential corpora move 0 of 1441 names, and the grid of MEASURED without the two words moves 0 parses, so every figure there stands. THE PASS IS ASKED ONLY WHERE A SUFFIX PIECE IS IN REACH: `_pieces.anchor_in_reach` walks back through the lone members in front of a declined member on tags alone and answers False wherever the pass must — the first piece past the members is no suffix piece, or there is none — and it is asked only before the pass exists, one lookup answering after that, so a run of members stays linear. An ordinary name ending in a declined member stops paying for the pass (frames, py3.11, before → after, e10e83b4 in brackets): `John Smith Ma` 250 → 246 (244), `Smith, John Ma` 285 → 279 (275), `Doe, Jane MA do` 374 → 369 (363), `Doe, John Q. Ma` 328 → 320 (316), `Jan vd Ma` 308 → 243 (238); a name the pass does read pays the test on top, `John Smith PhD MEng` 326 → 329 and `Doe, Jane PhD MEng` 359 → 362. On the no-comma peel the reach starts past the walk's leading piece, which never anchors — so a by-shape member, which reaches the pass only at the peel's second position (a lean of None with a word to spare is consumed first), finds nothing in reach, and the by-shape guard that stood there was deleted as unreachable. `segment_suffix_reading` calls `credential_anchors` (`first_kept=False`, a comma part's first piece being a run member like any other) in place of the inline copy NO MECHANISMS describes, and the keep-in-step note goes with it. The pass is built only where the reach test finds a suffix piece in front of a declined member, which no member opening the part has, and a part whose name word stands ahead of any member returns before one is asked: a plain comma name pays nothing (`Smith, John` 182, `Doe, John MA` 283, `Smith, J. Q.` 246, `Smith, Ed` and `Smith, Ma` 213, `Smith, Ed John` 270, all as before the call replaced the copy; built on every declined member instead, the last three paid 215, 215 and 273), a title in front pays the reach test alone (`Smith, Dr. Ma` 253 → 254), and a part the company decides pays three frames over the copy (`Smith, PhD MEng` 273 → 276, `Smith, PhD Ma` 271 → 274; `Smith, PhD Ma John` 330 → 335, its pass built for 'Ma' before 'John' ends the part). THE TWO-INPUT CHECK's second control reads 42, not 30, with no change to the property: the walk's leading piece is now held out of the pass twice on the no-comma peel, by `credential_anchors` and by the reach test, either alone suffices, and with both removed the by-shape heads the deleted guard held back (`PhD X.Y.Z.`) fail beside the listed ones; the first control is unmoved at 48. And the 2026-09-28 C2 bullet's "a TAIL segment, which assign reads wholly as suffixes" holds outside a maiden clause standing in it — `Jane Doe, PhD, Jr nee van Ma` keeps maiden 'van Ma' — which rules.md#C2's statement now says. The reach test's one-way exactness is held by tests/v2/pipeline/test_pieces.py's `test_anchor_in_reach_never_hides_an_anchor`, over the case table's texts and runs of one to three words behind two heads; its recorded negative control, the reach test reading a `vocab:suffix` token as no suffix piece, fails at 878 member positions. - 2026-09-27 (#544) — THE 2026-09-15 ACCEPTED ITEM (ii) IS REVERSED, and the bullet stands as it landed. `abdul Smith Jr Ma` reads given 'abdul', family 'Smith', suffix 'Jr Ma' — 2.3.0's reading — because the unambiguous 'Jr' in front of the Title-case 'Ma' anchors it (the entry above): the peel takes both, and P5's reserve, now seeing the family the join would take, declines the join. Item (i) stands. The case row is now `a_credential_in_front_anchors_a_declined_pick`, and rules.md#S2's Accepted block names the shapes that still keep the company out of reach. - 2026-09-27 (Derek), #544 — THE #531 PAIRING GAINS A SECOND EXCEPTION, and CAPITALS DECIDE FOR `do` above stands as it landed. That bullet let only a positive credential lean override P6's attachment at the given slot; an unambiguous credential IN FRONT of the member now overrides it as well, the degree being a second and stronger signal: `doe, jane v phd do` reads suffix 'v phd do' and reports `suffix-or-name` where it read family 'do doe' and reported P6's fork. The one-case record the pairing protects has nothing in front of its particle, so `NASCIMENTO, EDSON ARANTES DO` still reads family 'DO NASCIMENTO' and reports `particle-or-given`. +- 2026-09-30 (Derek), #563 — PAIRED INITIALS ARE THE EXCEPTION TO C1'S NAME-WORD COUNT, AND A FLIP ON THE DOTTED SHAPE ALONE IS SILENT. The 2026-09-14 (Derek) entry above, #516's DOTTED HALF IS A SWITCH, read `John Smith, A.B.` as suffix 'A.B.' on the count of two name words before the comma. The count cannot tell a given name and a family name from ONE surname written in two words, and the dotted shape's own most common member is a person's initials, so the unreleased tree read `García Márquez, G.J.` as given 'García', family 'Márquez', suffix 'G.J.', `De La Cruz, M.J.` with no given name at all, and `van der Berg, A.J.` as given 'van' — where 1.4.0 through 2.3.0 read each as initials. A particle check would have caught only the last two; `García Márquez` and `Lloyd Webber` look exactly like `John Smith`. What the parser CAN see is the token, so the line is drawn there: two single letters joined by a period, with or without one after the second (`_vocab.is_paired_initials`; `García Márquez, G.J` reads as `G.J.` does), read as the given name after the comma and report the fork, and only an UNAMBIGUOUS suffix word IN FRONT of them that is not also title vocabulary, or a second word the class admits by dotted shape in the same part, makes them the credential run (`John Smith, PhD X.Y.`, `John Smith, X.Y. P.Q.`). Three letters or more (`X.Y.Z.`) and any longer chunk (`B.Tech.`) read by the count as before. Three choices made with it (Derek, 2026-09-30): (i) FRONT ONLY — a word behind the pair does not speak, so `De La Cruz, M.J. PhD` keeps given 'M.J.' and `John Smith, A.B. PhD` reads given 'A.B.', the direction rules.md#S2's company clause already takes; a LISTED dotted word behind does not speak either (`John Smith, A.B. Ph.D.` → given 'A.B.'), the second-word exception being for words the class admits by shape; (ii) SILENT FLIP — a flip resting only on words the class admits by dotted shape reports nothing, `B.Tech.` included, because the only such word a reader takes for a name is a pair of initials, and the comma flips a pair only where a word beside it has already said otherwise; a LISTED member in the same part still reports the flip (`John Smith, PhD MEng`, `John Smith, X.Y.Z. MA`); (iii) THE COMMA SLOT ONLY — the trailing slot (`John Smith R.T.`), the given part's trailing slot after a family comma (`Doe, John R.T.`) and the maiden clause keep #516's reading and its report; #563 leaves the no-comma question open. A generational word in front speaks like any unambiguous suffix word (`García Márquez, Jr. G.J.` → suffix 'Jr. G.J.'), as S2's company clause has it — unless it is also a title, so `García Márquez, Sr G.J.` reads title 'Sr', given 'G.J.'. ACCEPTED: `García Márquez, G.J.R.` reads suffix 'G.J.R.' — initials are conventionally written apart, as separate words this rule never reads (`García Márquez, G. J. R.` gives given 'G.', middle 'J. R.'), and three letters run together are far likelier a credential. REVERSED BY THIS ENTRY rather than edited, for the dotted half at the first slot after a comma only: the #516 SWITCH entry's `John Smith, A.B.` → suffix and its "reports either way"; the 2026-09-14 NO NEW AmbiguityKind entry's "Every decision at these slots emits" SUFFIX_OR_NAME; and rules.md#C1's "A decision either way at this comma is reported". ALSO FIXED: the trailing slot's report described a by-shape pick in the listed member's words ("written without periods is both a post-nominal"), false of `X.Y.Z.` on both counts; it now says the word is shaped like a post-nominal but listed in no vocabulary. + REVIEW ROUND, same day. The docs review found the first version let a word of both the title and the suffix vocabulary speak for the pair, since C1 counts such a word opening the part as a suffix word: `García Márquez, Ms G.J.` read given 'García', family 'Márquez', suffix 'Ms G.J.' with NO report — #563's own defect, made silent — where 2.3.0 read title 'Ms', given 'G.J.'. The twelve title/suffix duals (`ms`, `md`, `sr`, `lt`, `sa`, `ra`, `vc`, `cpl`, `cpo`, `cpt`, `csm`, `sgm`) all reached it. Fixed by letting only a suffix word that is not title vocabulary speak: in front of a pair, a dual is the title of the given part the pair opens, as S2 already reads a dual in that part's title run, and `Ms G.J.`, `MD G.J.`, `Lt G.J.` now read title plus given as every release did. The title lookup runs only once a pair is met. The code review's frame count found the single-token path asking `ambiguous_class_member` a second time, +2 frames on every reporting comma name (`John Smith, MA`, 252 → 254, profiler call events); for a word past the candidate test LISTED is exactly "no period" (`ambiguous_class_member` declines any period), which is now asked inline. Measured after both fixes against master, same counter: `Smith, John` 183 and `John Smith, MA` 252 on both trees, `John Smith, X.Y.Z.` 295 → 286 (the report it no longer builds), and `García Márquez, G.J.` 338 → 350, the family-comma reading and its report costing more than the flip did. + SECOND REVIEW ROUND, same day, on the first round's fix commit. (a) TWO PAIRS (Derek): `De La Cruz, M.J. K.L.` read family 'De La Cruz', suffix 'M.J. K.L.' — no given name — in silence, the silent-flip rationale ("only paired initials are taken for a name") failing on exactly the case where each word speaking for a pair is itself a pair. The reading stands, two dotted groups not being how anyone writes a person's initials, but that flip now REPORTS: when the only words speaking for a pair are other pairs, every shape word in the part being one. `John Smith, X.Y. P.Q.` gains the report with it; `John Smith, X.Y.Z. G.J.` and `John Smith, PhD G.J. K.L.`, where something else speaks, stay silent. The same rule reports `García Márquez, Ms G.J. K.L.`, whose honorific the run takes, where the first round's fix took it in silence. (b) A CLASS MEMBER IN FRONT SPEAKS FOR NOTHING (Derek): `García Márquez, Ed G.J.`, `Ma G.J.` and `MA G.J.` flipped on the member in front, where S2's company lets only an unambiguous credential speak for a member. Now the comma keeps the family, given 'Ed' — and 'G.J.' then ends the given part, the slot (iii) left alone, so it reads suffix 'G.J.' as `Doe, John R.T.` does, both forks reported; 1.4.0 and 2.3.0 gave middle 'G.J.'. Resolving that slot is the no-comma question #563 leaves open. (c) WORDING, no behavior: C1's sentence counting a dual opening a part as a suffix word now defers to the pair's titles; the pair's report reaches it only where it opens the part, `Smith, Ms G.J.` having always been silent (the comma report's reach is the first post-comma piece, 2026-09-18); and the silence is stated as "no listed word takes part", which covers `García Márquez, PhD G.J.`, where "dotted-shape words alone" did not. + SIMPLIFY ROUND, same day, behavior-identical (0 diffs over 21,604 parses: every corpus name, every quoted string in tests/v2/cases.py and 115 composed `pre, post` probes, under four policies, comparing fields and ambiguity details against the pre-round commit c125f69b; the same harness finds 252 diffs against master). One finding was a cost, not a style point: the speaker test scanned every word in front of EACH pair for a non-title, so `John Smith, MD MD ... G.J. G.J. ...` cost duals × pairs `_normalize` calls (163 at 8 of each, 1,387 at 32, py3.11). Only the first pair's scan can change the answer, since every later pair has the same words in front and more, so it is asked once: 107 and 395. `tests/v2/test_benchmark.py::test_the_paired_initials_title_scan_does_not_cost_quadratically` guards the ratio and fails at c125f69b. The run loop also asks LISTED as "no period", as the single-token test does, and `flip_reports` is set once after the run decision rather than piecemeal. + MEASURED 2026-09-30 against master b98b26e3, every name in this branch's `tools/differential/corpus*.jsonl` parsed on both trees with `nameparser.__file__` asserted on each side: 11 of 1453 distinct names differ, every one of them a name this change's rules.md examples and case rows put in the corpus (the two-pair names `De La Cruz, M.J. K.L.` and `John Smith, X.Y. P.Q.` are not among them: after the second round they read and report exactly as master does). THE POPULATION THAT COULD MOVE is the shape's, and the corpus barely holds it: over master's 1441 distinct names, 13 have a pair opening the part after the first comma, and `John Smith, A.B.` is the only one behind two or more NAME words with an unlisted, non-CJK pair (`Smith Jr., A.B.` has one name word, `Kenneth Clarke Q.C., M.P.` and `Virginia G. Essandoh, J.D.` hold listed acronyms, the rest one word) — so it is the only mover over that corpus, and the count is evidence about the corpus rather than about the rule's reach. Recompute: check out the parent into a separate worktree, parse every corpus name in each tree under `PYTHONSAFEPATH=1` with the tree's root first on `sys.path`, and diff `as_dict()` plus the sorted ambiguity kinds; for the population, take each name whose text after its first comma opens with a token matching `[^\W\d_]\.[^\W\d_]\.?` whole. ### indic-honorifics — the renunciate class and the Indic honorific vocabulary (2026-09-06, #346/#344/#343) diff --git a/docs/design/rules.md b/docs/design/rules.md index c1356afa..077c8546 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -1072,7 +1072,12 @@ S2. Rationale: generational suffixes and credentials are recognized three-word name gives up its family name to the acronym, while a two-word name keeps it — there are no words to spare there, so the class is considered and declined and only the fork is - reported. + reported. The dotted shape does not take every promise of the + class with it: at the first slot after a comma, C1 reads paired + initials as the given name unless something speaks for them, and + makes a flip to the credential run in which no listed member takes + part in silence, unless paired initials speak only for each + other. "John Smith Jr." → suffix="Jr." "John Smith M.A." → suffix="M.A." "John Smith PhD" → suffix="PhD" @@ -1194,7 +1199,12 @@ S3. Rationale: credentials are often written run together with trailing the given part after one, the word ending a maiden marker's clause (M2), and the part before a SUFFIX comma. The part before a FAMILY comma never reports, the comma - having already named it the family. Case says nothing here — the + having already named it the family. C1 states two exceptions at + the first slot after a comma: paired initials ('M.J.') read as + the given name unless something in the part speaks for them, and + a flip to the credential run in which no listed member of the + class takes part is not reported, unless paired initials speak + only for each other. Case says nothing here — the periods are the evidence — and three shapes are outside it: a single trailing period is not this shape at all, a chunk that is not wholly alphabetic is no acronym letter, and a word carrying @@ -1660,7 +1670,23 @@ C1. Rationale: a credential run after the comma means the name is in credential run however that part is written, and only where the count leaves the word a name — one name word before the comma — is the case read, capitals in a mixed-case name making it the - credential there too (S2). The same count reads a part of two or + credential there too (S2). Paired initials are the exception to + the count: two single letters joined by a period, with or without + a period after the second ('M.J.', 'M.J'), the one dotted shape a + person's own initials are written in, read as the given name + after the comma, because two words before the comma may be one + surname written in two ('García Márquez') and the count cannot + tell them from a given name and a family name. Only an + unambiguous suffix word in front of them that is not also title + vocabulary, or another word the class admits by its dotted shape + standing in the same part, makes them the credential run. A word + of this class in front of them speaks for nothing, as S2's + company has it, and where nothing but words of both the suffix + and the title vocabulary stands in front of them, those are the + titles of the given part the initials open, as S2 reads them + there. Paired initials that are not the only shape word in the + part are a credential however little else speaks for them, since + no one writes a person's initials as two dotted groups. The same count reads a part of two or more words as the credential run when every word of it is a suffix word or a word of this class, at least one of them of this class, and none of them a single-letter roman numeral, in any @@ -1669,7 +1695,8 @@ C1. Rationale: a credential run after the comma means the name is in while a multi-letter numeral or generational word stands in a run like any suffix word ('John Smith, III Ma'). A word of both the title and the suffix vocabulary opening such a part counts as - a suffix word there, the name before the comma being complete. Where every word + a suffix word there, the name before the comma being complete, + except as the titles of paired initials, above. Where every word of this class in the part is a LISTED word written in capitals in a mixed-case name, the writing has already made each of them the credential (S2), and the part reads as the credential run on that @@ -1677,7 +1704,16 @@ C1. Rationale: a credential run after the comma means the name is in alone carries no such lean, so a part holding one is read by the count. A decision either way at this comma is reported; for a run of words the decision is the flip to the - credential run, reported once over the whole part. A run the + credential run, reported once over the whole part. A flip in which + no listed word of this class takes part is the exception and is + made in silence, since only paired initials among such words are + ever taken for a name, and the flip reaches them only where + something else in the part has said otherwise — unless what said + so is nothing but other paired initials: each of those is a word + a reader takes for a name, and that flip is reported. Paired + initials the comma keeps as the given name report where they open + the part; behind a title they do not, the comma's report reaching + only the first word after it ('Smith, Ms G.J.'). A run the count leaves in the listing form reports as S2 reads the words in it: a word of this class read as the credential because a credential in front speaks for it reports, and one its own @@ -1757,9 +1793,28 @@ C1. Rationale: a credential run after the comma means the name is in "John Smith, Ed Ma" → suffix="Ed Ma" "Jane Doe, MS LAc" → suffix="MS LAc" "Smith, PhD MEng" → family="Smith" · boundary - "John Smith, A.B." → suffix="A.B." - "John Smith, A.B." unlisted_dotted_suffixes-off → given="A.B." + "John Smith, X.Y.Z." → suffix="X.Y.Z." + "John Smith, X.Y.Z." → ambiguities=() + "John Smith, X.Y.Z." unlisted_dotted_suffixes-off → given="X.Y.Z." "Smith, A.B." → given="A.B." · boundary + "García Márquez, G.J." → given="G.J." + "García Márquez, G.J." → family="García Márquez" + "John Smith, A.B." → given="A.B." + "John Smith, A.B." → ambiguities=("suffix-or-name",) + "John Smith, X.Y. P.Q." → suffix="X.Y. P.Q." + "De La Cruz, M.J. K.L." → ambiguities=("suffix-or-name",) + "John Smith, PhD X.Y." → suffix="PhD X.Y." + "García Márquez, G.J" → given="G.J" + "De La Cruz, M.J. PhD" → given="M.J." · boundary + "García Márquez, Ms G.J." → given="G.J." · boundary + "García Márquez, Ed G.J." → given="Ed" · boundary + "John Smith, A.B. Ph.D." → given="A.B." · boundary + Accepted: three or more initials run together with periods behind + a surname of two words read as a credential. Initials are + conventionally written apart ('J. R. R.'), and a token of three or + more single letters run together is far likelier a credential than + a given name and a middle name written as one word. + "García Márquez, G.J.R." → suffix="G.J.R." Accepted: a word of both the title and the unambiguous suffix vocabulary reads as the postnominal after a family comma in every spelling, the honorific's too — position decides for the diff --git a/docs/release_log.rst b/docs/release_log.rst index d87a9eac..b00c54ee 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -16,11 +16,11 @@ Release Log - **Fix a credential run losing the acronyms in it that are also names: a degree in front speaks for the acronym behind it, and a run after a comma is read whole.** ``HumanName("John Smith, Ed Ma")`` gives first ``John``, last ``Smith``, suffix ``Ed Ma``, where 2.0 through 2.3 gave first ``Ed``, middle ``Ma``, last ``John Smith`` -- 1.4.0's reading, restored: with two or more name words before the comma, a part made only of suffix words and acronyms that are also names, and holding no one-letter roman numeral, is a credential run however it is written, as a lone one already was (``John Smith, MA``), and ``parse()`` reports the call wherever the writing left it open: a run whose every such acronym is written in capitals in a mixed-case name is the credential run without a report, the capitals having decided it (``John Smith, PhD MA``, ``John Smith, MS MA``). A one-letter numeral keeps a part out of that rule, but not out of the next one: a degree behind the letter still speaks for the acronyms after it and the part is read whole, so ``john smith, v phd ma`` gives first ``john``, last ``smith``, suffix ``v phd ma``, where 2.3 gave first ``v``, middle ``ma``, last ``john smith``. ``john smith, md ma`` and ``John Smith, Ms Ma`` move the same way, where 2.0 through 2.3 gave title ``md``/``Ms`` -- the second is the accepted cost, ``Ms`` read as the suffix word it also is, as ``John Smith, Ms`` alone already reads it -- while one name word before the comma keeps the listing form (``Smith, Ms Ma`` gives title ``Ms``, first ``Ma``). At the end of a name, an acronym standing behind an unambiguous credential is read as that credential's company whatever its case: ``John Smith PhD MEng`` and ``Doe, Jane PhD MEng`` give suffix ``PhD MEng``, the fields every release gave (``PhD, MEng`` through 2.2), now reported, and ``doe, jane v phd do`` gives suffix ``v phd do`` where 2.3.0 gave last ``do doe`` -- a degree in front outranks the particle reading, as capitals already did. Only a credential IN FRONT speaks: ``Wang Ma PhD`` keeps last ``Ma``. After a one-word family comma the part it speaks for reads wholly as credentials and the acronym it decided is reported: ``Smith, PhD Ma`` gives last ``Smith``, suffix ``PhD Ma``, where 2.3 gave first ``PhD``, middle ``Ma``. A title that is also a credential (``MD``, ``Ms``) opening that part stays a title and nothing in the part speaks, so ``Smith, MD PhD Ma`` keeps title ``MD``, first ``PhD``, middle ``Ma`` and ``Smith, Ms MD Ma`` title ``Ms MD``, first ``Ma``, as 2.3 read them. A listed acronym written in period-closed chunks is written with its periods, so ``Wang M.Eng.`` gives suffix ``M.Eng.``, as ``Wang M.A.`` does and as 2.0 through 2.3 did. See the #544 entry under ``S2`` in ``docs/design/decisions.md`` (closes #544) - - **New Policy field unlisted_dotted_suffixes, on by default: a dotted acronym nobody has listed is read by position.** ``HumanName("John Smith X.Y.Z.")`` gives suffix ``X.Y.Z.`` where every release gave last ``X.Y.Z.``, while ``Jack X.Y.Z.`` keeps its surname, the same words-to-spare rule a listed acronym takes -- and both readings are reported. Case is irrelevant here: the periods are the signal, so ``john smith x.y.z.`` reads the same way. Words the vocabulary does know are untouched (``M.A.``, ``Ph.D.``, ``A.B.C.``), a single trailing period is still not this shape (``John Smith Xyz.`` keeps last ``Xyz.``), and a dotted run at the FRONT of a name is untouched (``J.R.R. Tolkien``). One accident retires with it: a dotted word whose only vocabulary matches were SINGLE ASCII CHARACTERS -- the roman numerals the suffix list holds, and the lone digit ``2`` -- was reading as a generational suffix, so ``Jack X.Y.I.`` gives last ``X.Y.I.`` again, as 1.4.0 read it, while ``Msc.Ed.``, ``JD.CPA`` and ``Lt.Gov.`` are unchanged. The digit is why a dotted VERSION STRING moves with them and moves SILENTLY: ``John Smith 1.4.2`` gives last ``1.4.2`` where 2.3 gave suffix ``1.4.2``, and ``John Smith, 1.4.2`` gives first ``1.4.2``, last ``John Smith``. Such a token reports nothing at any policy -- it is no acronym either, the shape reading wanting every chunk alphabetic -- and a version string read as a credential was the same accident this retirement removes. That retirement is NOT behind this switch and stands either way -- setting it to ``False`` reads an unlisted dotted word as name material by position instead (``John Smith X.Y.Z.`` keeps last ``X.Y.Z.``), the pre-2.4 reading for THAT half alone. See the ``S2`` and ``suffix-acronym-collisions`` entries of ``docs/design/decisions.md`` (closes #516) + - **New Policy field unlisted_dotted_suffixes, on by default: a dotted acronym nobody has listed is read by position.** ``HumanName("John Smith X.Y.Z.")`` gives suffix ``X.Y.Z.`` where every release gave last ``X.Y.Z.``, while ``Jack X.Y.Z.`` keeps its surname, the same words-to-spare rule a listed acronym takes -- and both readings are reported. After a comma the count is of the words before it, and two dotted single letters are the exception: they are how a person's initials are written, and two words before a comma may be one surname, so ``García Márquez, G.J.`` keeps first ``G.J.`` and last ``García Márquez`` and reports the fork, unless an unambiguous post-nominal in front of the initials that is not also a title, or another unlisted dotted word beside them, says otherwise (``John Smith, PhD X.Y.`` gives suffix ``PhD X.Y.``, while ``García Márquez, Ms G.J.`` keeps title ``Ms``, first ``G.J.``). Three letters or more read by the count, so ``John Smith, X.Y.Z.`` gives suffix ``X.Y.Z.`` -- and so does ``García Márquez, G.J.R.``, the accepted cost of the line, since initials are conventionally written apart (``García Márquez, G. J. R.``), as separate words this rule does not read (#563). Case is irrelevant here: the periods are the signal, so ``john smith x.y.z.`` reads the same way. Words the vocabulary does know are untouched (``M.A.``, ``Ph.D.``, ``A.B.C.``), a single trailing period is still not this shape (``John Smith Xyz.`` keeps last ``Xyz.``), and a dotted run at the FRONT of a name is untouched (``J.R.R. Tolkien``). One accident retires with it: a dotted word whose only vocabulary matches were SINGLE ASCII CHARACTERS -- the roman numerals the suffix list holds, and the lone digit ``2`` -- was reading as a generational suffix, so ``Jack X.Y.I.`` gives last ``X.Y.I.`` again, as 1.4.0 read it, while ``Msc.Ed.``, ``JD.CPA`` and ``Lt.Gov.`` are unchanged. The digit is why a dotted VERSION STRING moves with them and moves SILENTLY: ``John Smith 1.4.2`` gives last ``1.4.2`` where 2.3 gave suffix ``1.4.2``, and ``John Smith, 1.4.2`` gives first ``1.4.2``, last ``John Smith``. Such a token reports nothing at any policy -- it is no acronym either, the shape reading wanting every chunk alphabetic -- and a version string read as a credential was the same accident this retirement removes. That retirement is NOT behind this switch and stands either way -- setting it to ``False`` reads an unlisted dotted word as name material by position instead (``John Smith X.Y.Z.`` keeps last ``X.Y.Z.``), the pre-2.4 reading for THAT half alone. See the ``S2`` and ``suffix-acronym-collisions`` entries of ``docs/design/decisions.md`` (closes #516) - **New Policy field unlisted_caps_suffixes, off by default: an opt-in reading for an unlisted all-caps credential.** It reaches the core parser only -- ``Parser(policy=Policy(unlisted_caps_suffixes=True))`` -- since the field has no v1 ``Constants`` manager. With it on, ``.parse("John Smith XYZ")`` gives given ``John``, last ``Smith``, suffix ``XYZ``, and ``.parse("John Smith, XYZ")`` gives the same three fields. It is off by default because an all-caps surname is a real writing convention that shape cannot separate from a credential: ``Jean DUPONT``, ``Minjun KIM`` and ``Jean Pierre DUPONT`` are surnames in French and Korean records, and the last of those gives given ``Jean``, last ``Pierre``, suffix ``DUPONT`` with the switch on. Off, nothing changes and nothing is reported -- 1.4.0's reading for that whole class. Neither of the two new fields reaches the v1 ``Constants`` API, as ``lenient_comma_suffixes`` does not: a ``HumanName`` tracks the parser's own DEFAULTS, so the dotted reading above (default on) reaches it while this one (default off) cannot be turned on from there. See the ``S2`` entry of ``docs/design/decisions.md`` (closes #516) - - **The comma's own decision about an ambiguous credential is now reported.** ``parse("Smith, MA").ambiguities`` names ``suffix-or-name``, and so does every other decision at the ambiguous credential class -- before or after a comma, in either direction, with no new ``AmbiguityKind`` (the family-comma attachment fork already reported this way, e.g. ``parse("Berg, Jan vd")``). One report per decision: ``Smith, Ma`` reports that the word was kept as the given name just as ``Smith, MA`` reports that it was taken as a credential. The reading a SURNAME PARTICLE swallows is reported too, which no release before this one did: ``John van der Berg Ma`` gives last ``van der Berg Ma`` and names ``suffix-or-name``, where the chain took a word the credential reading had considered. ONE report goes away, because a comma segment the parser reads as a credential run is no longer called unrecognized: ``Steven Hardman, MD, DO, DDS`` no longer reports ``comma-structure``, on its written case. That is the whole of the losses over the differential corpora -- ``John Smith, MD, R.A.I.`` is quieted on its shape by the same change, but it never reported at 2.3.0 either, having only carried the flag inside this release's own development. The other movement an upgrader sees is a SWAP rather than a loss: ``Jack X.Y.I.`` reported ``given-or-family`` at 2.3.0 and reports ``suffix-or-name`` here, the dotted retirement above having handed it to the ambiguous class. Everything else at this class is a GAIN, which is what the rest of this bullet describes. Two slots this bullet left silent no longer are, and the two bullets below close them: a credential trailing the GIVEN part of a family-comma listing now reads as a credential and reports either way, and so does one ending a maiden marker's clause. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` + - **The comma's own decision about an ambiguous credential is now reported.** ``parse("Smith, MA").ambiguities`` names ``suffix-or-name``, and so does every other decision at the ambiguous credential class -- before or after a comma, in either direction, with no new ``AmbiguityKind`` (the family-comma attachment fork already reported this way, e.g. ``parse("Berg, Jan vd")``). A flip of the comma in which no listed ambiguous acronym takes part is the exception and is made in silence: ``John Smith, X.Y.Z.`` and ``John Smith, PhD X.Y.`` report nothing, the only such word a reader takes for a name being a pair of initials, which the comma reads as the given name unless something beside it has already said otherwise. Two pairs speaking only for each other still make the credential run, and that flip reports: ``De La Cruz, M.J. K.L.`` gives last ``De La Cruz``, suffix ``M.J. K.L.`` (#563). One report per decision: ``Smith, Ma`` reports that the word was kept as the given name just as ``Smith, MA`` reports that it was taken as a credential. The reading a SURNAME PARTICLE swallows is reported too, which no release before this one did: ``John van der Berg Ma`` gives last ``van der Berg Ma`` and names ``suffix-or-name``, where the chain took a word the credential reading had considered. ONE report goes away, because a comma segment the parser reads as a credential run is no longer called unrecognized: ``Steven Hardman, MD, DO, DDS`` no longer reports ``comma-structure``, on its written case. That is the whole of the losses over the differential corpora -- ``John Smith, MD, R.A.I.`` is quieted on its shape by the same change, but it never reported at 2.3.0 either, having only carried the flag inside this release's own development. The other movement an upgrader sees is a SWAP rather than a loss: ``Jack X.Y.I.`` reported ``given-or-family`` at 2.3.0 and reports ``suffix-or-name`` here, the dotted retirement above having handed it to the ambiguous class. Everything else at this class is a GAIN, which is what the rest of this bullet describes. Two slots this bullet left silent no longer are, and the two bullets below close them: a credential trailing the GIVEN part of a family-comma listing now reads as a credential and reports either way, and so does one ending a maiden marker's clause. See the ``S2`` and ``C1`` entries of ``docs/design/decisions.md`` - **Fix a credential ending the given part of a family-comma listing being read as a middle name in silence.** ``HumanName("Doe, John MA")`` gives first ``John``, last ``Doe``, suffix ``MA``, where 2.0 through 2.3 gave middle ``MA`` -- and 1.4.0 gave the suffix, so this restores v1's reading for that half. The comma has already named the family and the first word after it is the given name, so the words-to-spare count that governs the comma-less form is satisfied by construction and the writing decides alone: ``Doe, John Ma`` keeps middle ``Ma``, written the way a name is written, and ``Doe, John Ed`` keeps middle ``Ed``. Either reading is now reported, and the report belongs to the SPELLING rather than to the fields -- a declined name re-rendered without its comma, ``John Ma Doe``, re-parses to those same three fields and reports nothing, the word no longer standing where the question is asked. A name word behind the credential still ends its reach and stays silent -- ``Doe, John MA Smith`` gives middle ``MA Smith`` and reports nothing -- while a credential run or a trailing title is transparent to it: ``Doe, John MA PhD`` gives suffix ``MA PhD`` and ``Doe, John MA Prof.`` gives title ``Prof.`` with suffix ``MA``. Two second-order movements an upgrader may see, both consequences of the word leaving the given part rather than of this rule reaching further: ``Doe, John Prof. MA`` now gives title ``Prof.`` where it gave middle ``Prof. MA``, the trailing-title chain reaching a word the credential used to hide; and ``Doe, John van MA`` gives last ``van Doe`` with suffix ``MA`` where it gave middle ``van MA``, the surname-particle rule reaching a particle the same way. A name written wholly in one case says nothing either way and takes the credential, which is what 1.4.0 read: ``DOE, MARY JO MA``, ``doe, john ma``, ``田中, 太郎 MA`` and ``김, 민준 MA`` all give a suffix. The unlisted dotted spelling moves with them without the parity claim -- ``Doe, John X.Y.Z.`` gives suffix ``X.Y.Z.`` where 1.4.0 and 2.3.0 both gave a middle name -- to match the comma-less ``John Doe X.Y.Z.``. One word is carved out: ``do`` is the only member of this class that is also a surname particle, so capitals decide it and, with nothing in front of it, the particle reading keeps every other spelling (a degree in front is the other exception, the next-but-one entry). ``Doe, John DO`` gives suffix ``DO``, while ``Doe, John do``, ``Doe, John Do``, ``DOE, JOHN DO`` and ``doe, john do`` are unchanged and keep the particle-or-given report they already had. In a name written wholly in one case the two cannot be told apart, so ``SMITH, JOHN DO`` keeps last ``DO SMITH`` as ``NASCIMENTO, EDSON ARANTES DO`` does -- right about the Portuguese record, wrong about the osteopath, and the report is how a caller finds the second. See the ``S2`` and ``P6`` entries of ``docs/design/decisions.md`` (closes #531) diff --git a/nameparser/_pipeline/_assign.py b/nameparser/_pipeline/_assign.py index afcd0784..7785fe88 100644 --- a/nameparser/_pipeline/_assign.py +++ b/nameparser/_pipeline/_assign.py @@ -83,7 +83,7 @@ WorkToken, _AMBIGUOUS_CREDENTIAL_TAGS, _NEVER_FLIPPED, copy_with, ) from nameparser._policy import Policy, Script -from nameparser._types import AmbiguityKind, Role +from nameparser._types import SHAPE_ACRONYM_TAG, AmbiguityKind, Role def _set_roles(tokens: list[WorkToken], piece: tuple[int, ...], role: Role) -> None: @@ -433,11 +433,18 @@ def _assign_main(seg_idx: int, state: ParseState, taken, declined = ( ("a suffix", "a name part") if token.role is Role.SUFFIX else (f"a {token.role.value} name", "a post-nominal")) + # A word in the class by SHAPE (rules.md#S2, #S3) is in no + # wordlist and may be written with its periods, so the listed + # member's wording would misdescribe it on both counts (#563) + what = ("is shaped like a post-nominal but listed in no " + "vocabulary, so it may be an ordinary name" + if SHAPE_ACRONYM_TAG in token.tags + else "written without periods is both a post-nominal " + "and an ordinary name") ambiguities.append(PendingAmbiguity( AmbiguityKind.SUFFIX_OR_NAME, - f"{token.text!r} written without periods is both a " - f"post-nominal and an ordinary name; read as {taken} " - f"rather than {declined}", + f"{token.text!r} {what}; read as {taken} rather than " + f"{declined}", piece)) # leading ambiguous particle read as a name (#121 surfaced) if name_pieces: diff --git a/nameparser/_pipeline/_segment.py b/nameparser/_pipeline/_segment.py index e4ed17bc..e47675e9 100644 --- a/nameparser/_pipeline/_segment.py +++ b/nameparser/_pipeline/_segment.py @@ -47,14 +47,16 @@ """ from __future__ import annotations +from nameparser._lexicon import _normalize from nameparser._pipeline._pieces import own_words from nameparser._pipeline._state import ( ParseState, PendingAmbiguity, Structure, comma_bucket, copy_with, ) from nameparser._pipeline._vocab import ( ambiguous_class_candidate, ambiguous_class_member, ambiguous_lean, - caps_shape_candidate, is_one_case, is_single_letter_numeral, - is_wholly_suffix, name_word_count, run_word_fold, + caps_shape_candidate, is_one_case, is_paired_initials, + is_single_letter_numeral, is_wholly_suffix, name_word_count, + run_word_fold, ) from nameparser._types import AmbiguityKind @@ -198,6 +200,27 @@ def class_run(seg: tuple[int, ...]) -> bool: and ambiguous_class_candidate( state.tokens[groups[1][0]].text, state.lexicon, state.policy)) + # Whether the flip reports: rules.md#C1, "A flip in which no listed + # word of this class takes part is the exception and is made in + # silence" -- set below by a listed word taking part, by the caps + # run, or (run test) by paired initials spoken for only by each + # other (#563). + flip_reports = False + if candidate: + lone = state.tokens[groups[1][0]].text + # For a word that passed the candidate test, LISTED is exactly + # "no period": `ambiguous_class_member` declines any period, + # and the dotted shape needs one. Asked inline -- a second + # membership call cost every reporting comma name ('John + # Smith, MA') two frames to learn what this already says. + flip_reports = "." not in lone + # rules.md#C1: "Paired initials are the exception to the + # count" -- 'García Márquez, G.J.' has two words before the + # comma and one surname. Alone in the part, nothing speaks + # for them, so the structure stays the family comma and + # assign reads and reports them as the given name (#563). + if not flip_reports and is_paired_initials(lone): + candidate = False # The all-caps half (Policy.unlisted_caps_suffixes, #516) is the # FIRST shape this class can wear across more than one token -- # 'LEED AP' is two separate all-caps words, not one glued acronym @@ -231,6 +254,7 @@ def class_run(seg: tuple[int, ...]) -> bool: one_case=False) for i in groups[1])): candidate = case_class() is False + flip_reports = candidate # rules.md#C1: "The same count reads a part of two or more words as # the credential run when every word of it is a suffix word or a # word of this class, at least one of them of this class" -- the @@ -268,6 +292,19 @@ def class_run(seg: tuple[int, ...]) -> bool: members: list[str] = [] rest: list[str] = [] settled = True + # #563, rules.md#C1: "Only an unambiguous suffix word in front + # of them that is not also title vocabulary, or another word + # the class admits by its dotted shape standing in the same + # part, makes them the credential run." `unspoken_pair`: the + # FIRST pair has only class members ('Ed G.J.') or titles ('Ms + # G.J.') in front of it. Only the first can change the answer + # -- every later pair has the same words in front and more -- + # so the title scan runs once, not once per pair, which made + # 'MD MD ... G.J. G.J. ...' quadratic. + unspoken_pair = False + any_listed = False + shaped = 0 + pairs = 0 lexicon = state.lexicon for i in groups[1]: text = state.tokens[i].text @@ -276,12 +313,23 @@ def class_run(seg: tuple[int, ...]) -> bool: or (fold == "ask" and ambiguous_class_candidate( text, lexicon, state.policy))) if is_member: + # LISTED is exactly "no period" for a member, as at the + # single-token test above + listed = "." not in text + if listed: + any_listed = True + else: + shaped += 1 + if is_paired_initials(text): + pairs += 1 + if pairs == 1 and all( + _normalize(w) in lexicon.titles + for w in rest): + unspoken_pair = True members.append(text) # the lean is the LISTED set's alone (S2): a member # admitted by shape ('X.Y.Z.') is read by the count - settled = (settled and text.isupper() - and (fold == "member" - or ambiguous_class_member(text, lexicon))) + settled = settled and listed and text.isupper() # "reject" first: a name word ends the run without the # numeral test's frame, and the numeral cannot be a # "reject" (it is suffix vocabulary, so it folds "defer") @@ -301,13 +349,20 @@ def class_run(seg: tuple[int, ...]) -> bool: # asked inline so a run with a Title-case member never # forces the case fact here; `ambiguous_lean` is what # answers. + # "unless what said so is nothing but other paired + # initials": every shape word a pair, the first unspoken. + # Alone, the pair keeps the family comma; with others, the + # run flips and reports. + pair_only = unspoken_pair and pairs == shaped candidate = (bool(members) + and not (pair_only and shaped == 1) and (not rest or is_wholly_suffix( rest, lexicon, state.policy)) and not (settled and case_class() is False and all(ambiguous_lean(t, False) == "credential" for t in members))) + flip_reports = candidate and (any_listed or pair_only) # Computed only where `candidate` is true, alongside `case_class()` # -- the same lazy gate: a non-candidate comma name never counts # its pre-comma words either. Hoisted to a local because the @@ -326,7 +381,7 @@ def class_run(seg: tuple[int, ...]) -> bool: or (pre_comma_names is not None and pre_comma_names >= 2)) else Structure.FAMILY_COMMA) ambiguities = list(state.ambiguities) - if candidate and structure is Structure.SUFFIX_COMMA: + if flip_reports and structure is Structure.SUFFIX_COMMA: # The first report of the comma's OWN structure call in the # library, and it is emitted for the branch taken HERE only -- # the flip. Where the structure did not move, that token's @@ -351,7 +406,7 @@ def class_run(seg: tuple[int, ...]) -> bool: # expression is its own frame on 3.11 regardless of element # count (unlike a list comprehension, which PEP 709 inlines # only from 3.12), so the join alone cost every REPORTING - # comma name (`John Smith, MA`, `John Smith, A.B.`, `Davis + # comma name (`John Smith, MA`, `John Smith, Ed`, `Davis # Royce, Ed`) +2 frames at the DEFAULT policy, a path this # switch must not touch at all (#516 review round, F4). # diff --git a/nameparser/_pipeline/_vocab.py b/nameparser/_pipeline/_vocab.py index 3b46e034..84fd6dcf 100644 --- a/nameparser/_pipeline/_vocab.py +++ b/nameparser/_pipeline/_vocab.py @@ -351,6 +351,23 @@ def _dotted(text: str) -> bool: or _CHUNKED.fullmatch(text) is not None) +# #563: exactly two single letters, each closed by a period -- the way +# a given and a middle initial are written run together ('M.J.', +# 'G.J.'). The trailing period is optional, as `period_joined_vocab`'s +# shape verdict lets it be. +_PAIRED_INITIALS = re.compile(r"[^\W\d_]\.[^\W\d_]\.?") + + +def is_paired_initials(text: str) -> bool: + """Two single letters run together with periods ('M.J.'): the one + unlisted dotted shape a person's own initials take, so after a + comma it is a name until something in front of it, or another + dotted word beside it, says it is a credential (rules.md#C1). + Three or more letters ('X.Y.Z.') and any longer chunk ('B.Tech.') + are outside it.""" + return _PAIRED_INITIALS.fullmatch(text) is not None + + def suffix_as_written(n: str, text: str, lexicon: Lexicon) -> bool: """Counts as a suffix as written, with NO initial veto (the veto differs by caller): unambiguous suffix vocabulary, or an ambiguous @@ -830,7 +847,7 @@ def is_wholly_suffix(texts: Sequence[str], lexicon: Lexicon, self-contradicting report ("holds 1 name words, so it is read as a credential run") -- proved by mutation testing to be otherwise unreached: nothing but this predicate's own two unit tests - depended on it, and 'John Smith, A.B.' still flips correctly + depended on it, and 'John Smith, X.Y.Z.' still flips correctly through `pre_comma_names >= 2` alone (#516 review round). Neither by-shape class, dotted or caps, reaches the comma form through this predicate at all -- only through `_vocab. diff --git a/tests/v2/cases.py b/tests/v2/cases.py index 687554d5..9444a053 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -2231,13 +2231,115 @@ def _check_cjk_shape_purity(self) -> None: "report is new", shape=1), Case("the_comma_structure_moves_with_the_shape_class", - "John Smith, A.B.", - {"given": "John", "family": "Smith", "suffix": "A.B."}, + "John Smith, X.Y.Z.", + {"given": "John", "family": "Smith", "suffix": "X.Y.Z."}, classification="fix(#516)", - ambiguities=("suffix-or-name",), notes="two NAME words before the comma read the part after " "it as the credential run, for the by-shape half " - "exactly as for the listed one (rules.md#C1)", + "exactly as for the listed one (rules.md#C1). Silent " + "since #563: three letters run together have no name " + "reading a reader would weigh", + shape=3), + Case("paired_initials_after_a_two_word_surname_stay_the_given", + "García Márquez, G.J.", + {"given": "G.J.", "family": "García Márquez"}, + ambiguities=("suffix-or-name",), + notes="#563: two words before the comma may be ONE surname, " + "and two dotted letters are how a person's own initials " + "are written, so the count gives way (rules.md#C1). " + "2.4's first reading was given 'García', family " + "'Márquez', suffix 'G.J.'; 1.4.0 and 2.3.0 read it as " + "here, and only the report is new", + shape=3), + Case("paired_initials_after_a_full_name_are_still_a_fork", + "John Smith, A.B.", + {"given": "A.B.", "family": "John Smith"}, + ambiguities=("suffix-or-name",), + notes="#563: the same shape as the row above with a given " + "name before the comma, which the parser cannot see. " + "A.B. is also the degree, so the report is the point: " + "the reading goes to the initials, and the fork is said", + shape=3), + Case("a_credential_behind_paired_initials_does_not_speak", + "De La Cruz, M.J. PhD", + {"given": "M.J.", "family": "De La Cruz", "suffix": "PhD"}, + ambiguities=("suffix-or-name",), + notes="#563, S2's company rule: only a credential IN FRONT " + "speaks for a word, so the degree behind the initials " + "leaves them the given name. Every release reads it " + "this way; the report is new", + shape=3), + Case("a_credential_in_front_speaks_for_paired_initials", + "John Smith, PhD X.Y.", + {"given": "John", "family": "Smith", "suffix": "PhD X.Y."}, + classification="fix(#563)", + notes="the contrast to the row above: the degree in front " + "makes the pair a credential, and a flip on a dotted " + "word admitted by shape is silent (rules.md#C1). 2.3.0 " + "read given 'PhD', middle 'X.Y.'", + shape=3), + Case("a_title_in_front_of_paired_initials_does_not_speak", + "García Márquez, Ms G.J.", + {"title": "Ms", "given": "G.J.", "family": "García Márquez"}, + notes="#563 review round: 'Ms' is both title and suffix " + "vocabulary, and in front of paired initials it is the " + "title of the given part they open (rules.md#C1, as S2 " + "reads a dual in that part's title run). The first fix " + "let it speak and split the surname in silence. Every " + "release reads it this way", + shape=3), + Case("a_credential_in_front_of_paired_initials_speaks", + "García Márquez, PhD G.J.", + {"given": "García", "family": "Márquez", "suffix": "PhD G.J."}, + classification="fix(#563)", + notes="the contrast to the row above: 'PhD' is no title, so " + "it speaks for the pair and the part is the credential " + "run, flipped in silence", + shape=3), + Case("a_second_dotted_word_speaks_for_paired_initials", + "John Smith, X.Y. P.Q.", + {"given": "John", "family": "Smith", "suffix": "X.Y. P.Q."}, + classification="fix(#563)", + ambiguities=("suffix-or-name",), + notes="nobody writes initials as two dotted groups, so a " + "second by-shape word makes the part the credential " + "run (rules.md#C1) -- reported, since here the pairs " + "speak only for each other and each is a word a reader " + "takes for a name. 1.4.0 read given 'X.Y.', middle " + "'P.Q.'", + shape=3), + Case("two_pairs_behind_a_two_word_surname_report_the_flip", + "De La Cruz, M.J. K.L.", + {"family": "De La Cruz", "suffix": "M.J. K.L."}, + classification="fix(#563)", + ambiguities=("suffix-or-name",), + notes="#563 second review round: the cost of the two-pair " + "line is a surname with no given name, so the flip it " + "makes is reported rather than silent (rules.md#C1). " + "1.4.0 read given 'M.J.', middle 'K.L.'", + shape=3), + Case("a_class_member_in_front_of_paired_initials_does_not_speak", + "García Márquez, Ed G.J.", + {"given": "Ed", "family": "García Márquez", "suffix": "G.J."}, + classification="fix(#516)", + ambiguities=("suffix-or-name", "suffix-or-name"), + notes="#563 second review round: 'Ed' is a listed ambiguous " + "member, and S2's company lets only an unambiguous " + "credential speak, so the comma keeps the family. 'G.J.' " + "then ends the given part, the slot #563 left alone, " + "where the dotted shape reads a credential as in 'Doe, " + "John R.T.'; both forks report. 1.4.0 and 2.3.0 read " + "middle 'G.J.'", + shape=3), + Case("three_run_together_initials_behind_a_double_surname_read_as_a_credential", + "García Márquez, G.J.R.", + {"given": "García", "family": "Márquez", "suffix": "G.J.R."}, + classification="fix(#516)", + notes="the ACCEPTED cost of #563's line (rules.md#C1): " + "initials are conventionally written apart ('J. R. " + "R.'), so three letters run together read as a " + "credential even where the name before the comma is " + "one surname. 1.4.0 read given 'G.J.R.'", shape=3), Case("one_word_before_the_comma_keeps_the_given", "Smith, A.B.", {"given": "A.B.", "family": "Smith"}, diff --git a/tests/v2/pipeline/test_assign.py b/tests/v2/pipeline/test_assign.py index 0dbc10ab..50af5ebd 100644 --- a/tests/v2/pipeline/test_assign.py +++ b/tests/v2/pipeline/test_assign.py @@ -133,6 +133,25 @@ def test_the_family_comma_report_detail_is_verbatim() -> None: "a credential") +def test_a_by_shape_pick_is_not_described_as_a_listed_member() -> None: + # #563: the trailing slot's listed-member wording ("written without + # periods is both a post-nominal") is false of a word the class + # admits by SHAPE -- 'X.Y.Z.' is written WITH its periods and is in + # no wordlist. Both readings of the fork, so neither branch of the + # template can keep the old wording; 'MA' in the test below is the + # listed control that keeps it. + for text, taken, declined in ( + ("John Smith X.Y.Z.", "a suffix", "a name part"), + ("Jack X.Y.Z.", "a family name", "a post-nominal")): + out = _assigned(text, lexicon=Lexicon.default()) + (amb,) = [a for a in out.ambiguities + if a.kind is AmbiguityKind.SUFFIX_OR_NAME] + assert amb.detail == ( + f"'X.Y.Z.' is shaped like a post-nominal but listed in no " + f"vocabulary, so it may be an ordinary name; read as " + f"{taken} rather than {declined}"), text + + def test_jack_ma_s_two_detail_strings_are_verbatim() -> None: # The trailing slot's OWN report existed before #289 (a bare # ambiguous acronym was always a coin-flip); what #289 changes is diff --git a/tests/v2/pipeline/test_segment.py b/tests/v2/pipeline/test_segment.py index b4a212ed..836bbb3b 100644 --- a/tests/v2/pipeline/test_segment.py +++ b/tests/v2/pipeline/test_segment.py @@ -143,15 +143,60 @@ def test_structure_flips_for_the_ambiguous_class_on_a_name_word_count() -> None: def test_structure_flips_for_a_by_shape_member_too() -> None: # #516: an unlisted dotted token joins the class the same way, via - # `_vocab.ambiguous_class_candidate` -- 'A.B.' is two unclaimed + # `_vocab.ambiguous_class_candidate` -- 'X.Y.Z.' is three unclaimed # single-letter chunks, not vocabulary at all, so this is the - # by-shape twin of the test above. 'Smith Jr., A.B.' does not flip - # for the SAME reason 'Smith Jr., MA' does not (#516 review round: - # is_wholly_suffix must never admit the shape class here, or C1's - # legacy TOKEN-count disjunct flips it wrongly on 'Jr.'). - assert _segmented("John Smith, A.B.").structure is Structure.SUFFIX_COMMA - assert _segmented("Smith, A.B.").structure is Structure.FAMILY_COMMA - assert _segmented("Smith Jr., A.B.").structure is Structure.FAMILY_COMMA + # by-shape twin of the test above. 'Smith Jr., X.Y.Z.' does not + # flip for the SAME reason 'Smith Jr., MA' does not (#516 review + # round: is_wholly_suffix must never admit the shape class here, or + # C1's legacy TOKEN-count disjunct flips it wrongly on 'Jr.'). + assert _segmented("John Smith, X.Y.Z.").structure \ + is Structure.SUFFIX_COMMA + assert _segmented("Smith, X.Y.Z.").structure is Structure.FAMILY_COMMA + assert _segmented("Smith Jr., X.Y.Z.").structure \ + is Structure.FAMILY_COMMA + + +def test_paired_initials_need_a_word_to_speak_for_them() -> None: + # rules.md#C1 (#563): two dotted single letters are how a person's + # own initials are written, and two words before the comma may be + # one surname, so alone they keep the family comma. Each contrast + # pair differs in the one thing that speaks: a third letter, a + # credential IN FRONT (not behind), a second by-shape word. A + # generational word in front speaks like any suffix word; the + # title-vocabulary exception needs titles this lexicon lacks, so + # its contrast is the case-table pair + # a_title_in_front_of_paired_initials_does_not_speak. + for alone, spoken in (("García Márquez, G.J.", "García Márquez, G.J.R."), + ("De La Cruz, M.J. PhD", "De La Cruz, PhD M.J."), + ("John Smith, X.Y. MA", "John Smith, X.Y. P.Q."), + ("García Márquez, G.J.", "García Márquez, Jr G.J."), + # a class member in front speaks for nothing, + # as S2's company has it; an unambiguous one does + ("García Márquez, Ma G.J.", "García Márquez, PhD G.J.")): + assert _segmented(alone).structure is Structure.FAMILY_COMMA, alone + assert _segmented(spoken).structure is Structure.SUFFIX_COMMA, spoken + + +def test_a_flip_no_listed_member_takes_part_in_is_silent() -> None: + # rules.md#C1 (#563): a flip in which no listed member takes part + # is silent, while a LISTED member in the same position still + # reports the flip -- the control that keeps the silence from + # being a lost emitter -- and so do paired initials spoken for + # only by each other, each of them a word a reader takes for a + # name. A pair beside a three-letter word, or behind PhD, is + # spoken for by something else and stays silent. + for text in ("John Smith, X.Y.Z.", "John Smith, B.Tech.", + "John Smith, X.Y.Z. P.D.Q.", "John Smith, PhD P.D.Q.", + "John Smith, X.Y.Z. G.J.", "John Smith, PhD G.J. K.L."): + out = _segmented(text) + assert out.structure is Structure.SUFFIX_COMMA, text + assert _flip_reports(out) == [], text + for text in ("John Smith, Ma", "John Smith, PhD Ma", + "John Smith, X.Y.Z. MA", "De La Cruz, M.J. K.L.", + "John Smith, X.Y. P.Q."): + out = _segmented(text) + assert out.structure is Structure.SUFFIX_COMMA, text + assert _flip_reports(out) == [_texts(out, out.segments[1])], text def test_segment_records_the_case_fact_only_where_it_asked() -> None: @@ -222,7 +267,7 @@ def test_the_name_word_count_reads_a_run_as_it_reads_one_word() -> None: for text in ("John Smith, PhD Ma", "John Smith, Ed Ma", "John Smith, Ma PhD", "john smith, phd ma", "JOHN SMITH, PHD MA", "John Smith, MD Ma", - "John Smith, A.B. PhD", "John Smith, Ph. D. Ma"): + "John Smith, X.Y.Z. MA", "John Smith, Ph. D. Ma"): out = _segmented(text) assert out.structure is Structure.SUFFIX_COMMA, text assert _flip_reports(out) == [_texts(out, out.segments[1])], text diff --git a/tests/v2/test_benchmark.py b/tests/v2/test_benchmark.py index dcf82175..0fd38a25 100644 --- a/tests/v2/test_benchmark.py +++ b/tests/v2/test_benchmark.py @@ -687,6 +687,51 @@ def test_a_clause_link_run_does_not_cost_quadratically() -> None: f"again (#397)") +# #563 simplify round: segment's paired-initials speaker test scanned +# every word in front of EACH pair for a non-title speaker, so a comma +# run of title/suffix duals followed by pairs ('MD MD ... G.J. G.J. +# ...') cost duals x pairs `_normalize` calls. Only the first pair's +# scan can change the answer, and the fix asks it once. A `_SHAPES` row +# cannot express it -- the run needs the 'John Smith, ' prefix -- and +# the cost is Python-level, so it is counted in `_normalize` frames, +# which isolates the scan from the rest of the parse. Measured +# 2026-09-30 on py3.11 through `_frames_for(..., only="_normalize")`, +# k duals and k pairs: 107 at k=8 and 395 at k=32 on this tree (3.7x), +# against 163 and 1,387 at c125f69b (8.5x), where every pair rescanned. +# 6.0 sits between them; counts are deterministic, so the margins are +# for future shape changes, not noise. +_PAIR_SCAN_SMALL = 8 +_PAIR_SCAN_LARGE = 32 +_PAIR_SCAN_MAX_RATIO = 6.0 + + +def test_the_paired_initials_title_scan_does_not_cost_quadratically() -> None: + if sys.getprofile() is not None: + pytest.skip("a profile hook is already installed; this test owns it") + small_text = "John Smith, " + "MD " * _PAIR_SCAN_SMALL + "G.J. " * _PAIR_SCAN_SMALL + large_text = "John Smith, " + "MD " * _PAIR_SCAN_LARGE + "G.J. " * _PAIR_SCAN_LARGE + # REACHABILITY: the flip REPORTS only where the scan ran and found + # nothing but titles in front of the first pair, every shape word + # being a pair (rules.md#C1's "nothing but other paired initials"). + # A change that stopped the scan being asked -- 'MD' leaving the + # titles, the pair test moving -- loses the report and fails here + # rather than leaving this guard measuring nothing. + for text, k in ((small_text, _PAIR_SCAN_SMALL), + (large_text, _PAIR_SCAN_LARGE)): + name = parse(text) + assert len(name.suffix.split()) == 2 * k, text + assert [a.kind.value for a in name.ambiguities] == ["suffix-or-name"] + small = _frames_for(small_text, only="_normalize") + large = _frames_for(large_text, only="_normalize") + ratio = large / small + assert ratio < _PAIR_SCAN_MAX_RATIO, ( + f"{_PAIR_SCAN_SMALL} duals and pairs cost {small} _normalize calls " + f"and {_PAIR_SCAN_LARGE} cost {large} -- {ratio:.1f}x for 4x the " + f"input, where this tree measures 3.7x and the per-pair rescan at " + f"c125f69b measured 8.5x. _segment.py's paired-initials title scan " + f"is running once per pair again (#563)") + + # THE ABSOLUTE COST OF A LINK, which the ratio above cannot see: a # change costing ONE MORE FRAME PER LINK moves both ends of the pair # and leaves the ratio where it was. Re-splitting the #397 follow-up's diff --git a/tests/v2/test_ledger_guards.py b/tests/v2/test_ledger_guards.py index 1350256c..22603096 100644 --- a/tests/v2/test_ledger_guards.py +++ b/tests/v2/test_ledger_guards.py @@ -1401,7 +1401,20 @@ def test_case_shape_ids_exist_in_the_inventory() -> None: # earlier and so cannot exercise this one); and the # alphabetic-chunk gate at a COMMA ('John Smith, 1.4', the # comma twin of the bare 'John Smith 1.4' already here). - "John Smith 田.中.", "Bridge (A.B)", "John Smith, 1.4"), + "John Smith 田.中.", "Bridge (A.B)", "John Smith, 1.4", + # 2026-09-30, #563: paired initials after a comma are the + # paired-initials rule's, never this one's + "John Smith, A.B.", "García Márquez, G.J."), + # #563: the shapes the exception does not reach. One word before + # the comma never flips (#516's reading, reported there); three + # letters run together read by the count (#516's rule claims them); + # a pair after a suffix-bearing family keeps the given by the + # name-word count alone; spaced initials are separate words. + "fix(#563) paired initials after a comma read as the given name unless a word speaks for them": + ("Smith, A.B.", "Tolkien, J.R.R.", "John Smith, X.Y.Z.", + "García Márquez, G.J.R.", "Smith Jr., A.B.", "Smith, J. R. R.", + "John Smith X.Y.", "García Márquez, Ms G.J.", + "García Márquez, Ed G.J."), # Policy.unlisted_caps_suffixes is OFF by default, so its whole # population is a probe here: the corpora run at the default, and # a rule of this arc reaching one of these names would mean the @@ -2447,11 +2460,22 @@ class _LatinCopy(NamedTuple): "John Smith, Ma", "John de Ma", "John van der Berg Ma", r"Smith Jr\., MA", r"Smith Jr\., Ma", "Smith, MA", "abdul Smith Berg Ma", "abdul Smith Ma", "john smith, ma"}), - frozenset({r"Jack X\.Y\.I\.", r"John Smith B\.Tech\.", - r"John Smith C\.H\.A\.", r"John Smith E\.S\.Q\.", - r"John Smith Q\.W\.E\.R\.T\.", r"John Smith X\.Y\.Z\.", - r"John Smith, A\.B\.", r"Smith, E\.S\.Q\.", - r"john smith x\.y\.z\."}), + frozenset({r"García Márquez, G\.J\.R\.", r"Jack X\.Y\.I\.", + r"John Smith B\.Tech\.", r"John Smith C\.H\.A\.", + r"John Smith E\.S\.Q\.", r"John Smith Q\.W\.E\.R\.T\.", + r"John Smith X\.Y\.Z\.", r"John Smith, X\.Y\.Z\.", + r"Smith, E\.S\.Q\.", r"john smith x\.y\.z\."}), + # #563's paired-initials rule, literal-anchored the same way and + # for the same reason: what selects its names is two dotted single + # letters and what stands around them, which no wordlist spells. + # The 1.4.0 ledger lists the three role movers alone. + frozenset({r"De La Cruz, M\.J\. K\.L\.", r"García Márquez, PhD G\.J\.", + r"John Smith, PhD X\.Y\.", r"John Smith, X\.Y\. P\.Q\."}), + frozenset({r"De La Cruz, M\.J\. K\.L\.", + r"De La Cruz, M\.J\. PhD", r"García Márquez, G\.J", + r"García Márquez, G\.J\.", r"García Márquez, PhD G\.J\.", + r"John Smith, A\.B\.", r"John Smith, A\.B\. Ph\.D\.", + r"John Smith, PhD X\.Y\.", r"John Smith, X\.Y\. P\.Q\."}), frozenset({"Doe, MA Smith", r"J\.A\. K\.D\.", r"Jack X\.Y\.Z\.", r"John Smith J\.u\.n\.i\.o\.r\.", r"John Smith R\.A\.I\.", "Royce, Ed", r"Smith Jr\., A\.B\.", r"Smith, A\.B\.", @@ -2486,7 +2510,14 @@ class _LatinCopy(NamedTuple): "Doe, John MA", "Doe, John MA JD", "Doe, John MA Jr", "Doe, John MA PhD", "Doe, John PhD MA", r"Doe, John Q\. MA", r"Doe, John X\.Y\.Z\.", + r"García Márquez, Ed G\.J\.", "Smith nee Jones, Jane MA", "doe, john ma"}), + # 2026-09-30, #563: 'García Márquez, Ed G.J.' joins that set in + # every 2.x ledger -- the same slot, a by-shape member ending the + # given part, reached because a class member in front of a pair + # no longer flips the comma -- and the 1.4.0 ledger's + # single-name dotted rule becomes this two-name alternation. + frozenset({r"Doe, John X\.Y\.Z\.", r"García Márquez, Ed G\.J\."}), frozenset({"Doe, John Ed", "Doe, John MA Ma", "Doe, John Ma", "Doe, Mary Jo Ma"}), # 2026-09-28: the 2.2.0 and 2.3.0 copies of the declining rule @@ -3712,8 +3743,11 @@ def _claim(rule: dict) -> _Claim: # 2026-09-28, #544: 392 -> 396; gains 'John Smith, PhD Ma', # 'Smith, MD PhD Ma', 'Smith, Ms MD Ma', 'Smith, PhD Ma'. # Reach, verified name by name. + # 2026-09-30, #563: 396 -> 408, the twelve names #563's rules.md#C1 + # examples and case rows add, every one written with a comma. + # Reach, verified name by name. "fix(comma-family) lone post-comma piece routes to suffix/title, not first": - _Claim(396, ('given', 'suffix', 'title'), "1c2cfbb3f881", None), + _Claim(408, ('given', 'suffix', 'title'), "6f26a85296fb", None), "fix(comma-family) a comma followed only by titles keeps the given/family split": _Claim(2, ('family', 'given'), "5bd9c6d96c38", None), "fix(comma-family) a comma followed only by titles keeps the given/family split, the C1 example": @@ -3802,8 +3836,11 @@ def _claim(rule: dict) -> _Claim: # 'Doe, Jane nee Smith PhD MEng', 'Jane Doe, MS LAc', 'John # Smith, Ed Ma' and 8 more. # 2026-09-28, #544: 392 -> 396; the same four comma names. + # 2026-09-30, #563: 396 -> 408, the twelve names #563's rules.md#C1 + # examples and case rows add, every one written with a comma. + # Reach, verified name by name. "fix(comma-precomma-family) pre-comma run reads as family, not given": - _Claim(396, ('family', 'given'), "1c2cfbb3f881", None), + _Claim(408, ('family', 'given'), "6f26a85296fb", None), # 2026-09-20, #397: retitled in place, reach and digest # unchanged -- the rule keeps 'Carod i', which the landing # leaves byte-identical. @@ -4083,8 +4120,11 @@ def _claim(rule: dict) -> _Claim: # name and neither widened. Verified name by name. # 2026-09-26, #535: 112 -> 113, 'Jane van der Berg nee Smith # Prof.' -- another particle chain, 'van der Berg'. Reach again. + # 2026-09-30, #563: 115 -> 117, 'De La Cruz, M.J. PhD' and + # 'De La Cruz, M.J. K.L.' -- a particle chain, 'De La Cruz'. + # Reach, verified name by name. "fix(initials-per-word) a particle chain inside a name part initials each word (facade, since 2.0.0)": - _Claim(115, ('_initials',), '5f9056683f2f', ('DEFAULT',)), + _Claim(117, ('_initials',), '0820bd80678d', ('DEFAULT',)), # 2026-09-23, #459: 18 -> 19, 'john smith ph. d.', rules.md#R4's # two-token line. Reach, verified name by name. "fix(initials-per-word) the Ph. D. merge initials each word (facade, since 2.0.0)": @@ -4182,8 +4222,17 @@ def _claim(rule: dict) -> _Claim: # `orders` DEFAULT. Same reasoning as the rule above: the # class is a shape the vocabulary does not spell, so the # probes and not this number are what bound it. + # 2026-09-30, #563: 9 -> 10, a different set. 'John Smith, A.B.' + # leaves for the #563 rule (paired initials, rules.md#C1), and + # 'John Smith, X.Y.Z.' and 'García Márquez, G.J.R.' join, the + # comma form's by-count movers. Literal-anchored, so the reach + # IS the mover list. "fix(#516) an unlisted dotted acronym is read by position": - _Claim(9, ('family', 'given', 'middle', 'suffix'), "9bbaf4e84dc0", ('DEFAULT',)), + _Claim(10, ('family', 'given', 'middle', 'suffix'), "9efed4efb90b", ('DEFAULT',)), + # 2026-09-30, #563: new. Literal-anchored to its movers -- + # the four role movers; a report-only diff is invisible at this baseline. + "fix(#563) paired initials after a comma read as the given name unless a word speaks for them": + _Claim(4, ('family', 'given', 'middle', 'suffix', 'title'), "ff266bb1798d", ('DEFAULT',)), # 2026-09-18, verification round. ONE name, and a rule this PR # did not earn: 'Doe, John MA' has read middle since 2.0.0 and # reads suffix at 1.4.0, so the diff exists at this baseline @@ -4213,8 +4262,11 @@ def _claim(rule: dict) -> _Claim: # trailing slot at all; and the fix(#380) decision read over # the two other collision words, five names this PR measured # byte-identical at cc78c960. + # 2026-09-30, #563: 1 -> 2, 'García Márquez, Ed G.J.' -- a + # class member in front of a pair keeps the family comma, and + # the pair then ends the given part. Literal-anchored. "fix(#531) a dotted credential ending the given part leaves 1.4's middle name": - _Claim(1, ('middle', 'suffix'), "5800f141483d", ('DEFAULT',)), + _Claim(2, ('middle', 'suffix'), "0d423d266341", ('DEFAULT',)), "fix(comma-family) an interior credential acronym stays a middle-name word": _Claim(1, ('middle', 'suffix'), "4831097f5067", ('DEFAULT',)), # 2026-09-19, #531 fix round: two rules land, each literal- @@ -4869,8 +4921,17 @@ def _claim(rule: dict) -> _Claim: # `orders` DEFAULT. Same reasoning as the rule above: the # class is a shape the vocabulary does not spell, so the # probes and not this number are what bound it. + # 2026-09-30, #563: 9 -> 10, a different set. 'John Smith, A.B.' + # leaves for the #563 rule (paired initials, rules.md#C1), and + # 'John Smith, X.Y.Z.' and 'García Márquez, G.J.R.' join, the + # comma form's by-count movers. Literal-anchored, so the reach + # IS the mover list. "fix(#516) an unlisted dotted acronym is read by position": - _Claim(9, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "9bbaf4e84dc0", ('DEFAULT',)), + _Claim(10, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "9efed4efb90b", ('DEFAULT',)), + # 2026-09-30, #563: new. Literal-anchored to its movers -- + # four role movers and five names whose only diff is the report. + "fix(#563) paired initials after a comma read as the given name unless a word speaks for them": + _Claim(9, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "3550ea3b5804", ('DEFAULT',)), # The report-only rule. `_ambiguities` alone, so a widening that # took a ROLE would change the roles here before it reached # the gate -- which is the one thing this row can say about a @@ -4896,8 +4957,11 @@ def _claim(rule: dict) -> _Claim: # `Doe, John DO` moves {middle, suffix} here and # {family, suffix} from 2.2.0 on, which is P6's own shipping # date showing through. + # 2026-09-30, #563: 15 -> 16, 'García Márquez, Ed G.J.' -- a + # class member in front of a pair keeps the family comma, and + # the pair then ends the given part. Literal-anchored. "fix(#531) a credential ending the given part of a family-comma listing reads as a credential": - _Claim(15, ('_ambiguities', 'middle', 'suffix'), "f17532bb4ff1", ('DEFAULT',)), + _Claim(16, ('_ambiguities', 'middle', 'suffix'), "2da482a17f0e", ('DEFAULT',)), "fix(#531) a member the writing declines keeps its name reading and reports the fork": _Claim(4, ('_ambiguities',), "c6d26d145ee1", ('DEFAULT',)), "fix(#531) capitals take the do collision from the family-comma particle attachment": @@ -5350,8 +5414,17 @@ def _claim(rule: dict) -> _Claim: # `orders` DEFAULT. Same reasoning as the rule above: the # class is a shape the vocabulary does not spell, so the # probes and not this number are what bound it. + # 2026-09-30, #563: 9 -> 10, a different set. 'John Smith, A.B.' + # leaves for the #563 rule (paired initials, rules.md#C1), and + # 'John Smith, X.Y.Z.' and 'García Márquez, G.J.R.' join, the + # comma form's by-count movers. Literal-anchored, so the reach + # IS the mover list. "fix(#516) an unlisted dotted acronym is read by position": - _Claim(9, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "9bbaf4e84dc0", ('DEFAULT',)), + _Claim(10, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "9efed4efb90b", ('DEFAULT',)), + # 2026-09-30, #563: new. Literal-anchored to its movers -- + # four role movers and five names whose only diff is the report. + "fix(#563) paired initials after a comma read as the given name unless a word speaks for them": + _Claim(9, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "3550ea3b5804", ('DEFAULT',)), # The report-only rule. `_ambiguities` alone, so a widening that # took a ROLE would change the roles here before it reached # the gate -- which is the one thing this row can say about a @@ -5377,8 +5450,11 @@ def _claim(rule: dict) -> _Claim: # `Doe, John DO` moves {family, suffix} here, P6 having # shipped in 2.3; at 2.0.0 and 2.1.0 the same name moves # {middle, suffix} instead. + # 2026-09-30, #563: 15 -> 16, 'García Márquez, Ed G.J.' -- a + # class member in front of a pair keeps the family comma, and + # the pair then ends the given part. Literal-anchored. "fix(#531) a credential ending the given part of a family-comma listing reads as a credential": - _Claim(15, ('_ambiguities', 'middle', 'suffix'), "f17532bb4ff1", ('DEFAULT',)), + _Claim(16, ('_ambiguities', 'middle', 'suffix'), "2da482a17f0e", ('DEFAULT',)), # 2026-09-28, #544: 4 -> 5; gains 'Smith, MD PhD Ma'. "fix(#531) a member the writing declines keeps its name reading and reports the fork": _Claim(5, ('_ambiguities',), "f73bc2fc5408", ('DEFAULT',)), @@ -5966,8 +6042,17 @@ def _claim(rule: dict) -> _Claim: # `orders` DEFAULT. Same reasoning as the rule above: the # class is a shape the vocabulary does not spell, so the # probes and not this number are what bound it. + # 2026-09-30, #563: 9 -> 10, a different set. 'John Smith, A.B.' + # leaves for the #563 rule (paired initials, rules.md#C1), and + # 'John Smith, X.Y.Z.' and 'García Márquez, G.J.R.' join, the + # comma form's by-count movers. Literal-anchored, so the reach + # IS the mover list. "fix(#516) an unlisted dotted acronym is read by position": - _Claim(9, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "9bbaf4e84dc0", ('DEFAULT',)), + _Claim(10, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "9efed4efb90b", ('DEFAULT',)), + # 2026-09-30, #563: new. Literal-anchored to its movers -- + # four role movers and five names whose only diff is the report. + "fix(#563) paired initials after a comma read as the given name unless a word speaks for them": + _Claim(9, ('_ambiguities', 'family', 'given', 'middle', 'suffix', 'title'), "3550ea3b5804", ('DEFAULT',)), # The report-only rule. `_ambiguities` alone, so a widening that # took a ROLE would change the roles here before it reached # the gate -- which is the one thing this row can say about a @@ -5993,8 +6078,11 @@ def _claim(rule: dict) -> _Claim: # `Doe, John DO` moves {middle, suffix} here and # {family, suffix} from 2.2.0 on, which is P6's own shipping # date showing through. + # 2026-09-30, #563: 15 -> 16, 'García Márquez, Ed G.J.' -- a + # class member in front of a pair keeps the family comma, and + # the pair then ends the given part. Literal-anchored. "fix(#531) a credential ending the given part of a family-comma listing reads as a credential": - _Claim(15, ('_ambiguities', 'middle', 'suffix'), "f17532bb4ff1", ('DEFAULT',)), + _Claim(16, ('_ambiguities', 'middle', 'suffix'), "2da482a17f0e", ('DEFAULT',)), "fix(#531) a member the writing declines keeps its name reading and reports the fork": _Claim(4, ('_ambiguities',), "c6d26d145ee1", ('DEFAULT',)), "fix(#531) capitals take the do collision from the family-comma particle attachment": @@ -6313,8 +6401,17 @@ def _claim(rule: dict) -> _Claim: # `orders` DEFAULT. Same reasoning as the rule above: the # class is a shape the vocabulary does not spell, so the # probes and not this number are what bound it. + # 2026-09-30, #563: 9 -> 10, a different set. 'John Smith, A.B.' + # leaves for the #563 rule (paired initials, rules.md#C1), and + # 'John Smith, X.Y.Z.' and 'García Márquez, G.J.R.' join, the + # comma form's by-count movers. Literal-anchored, so the reach + # IS the mover list. "fix(#516) an unlisted dotted acronym is read by position": - _Claim(9, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "9bbaf4e84dc0", ('DEFAULT',)), + _Claim(10, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "9efed4efb90b", ('DEFAULT',)), + # 2026-09-30, #563: new. Literal-anchored to its movers -- + # four role movers and five names whose only diff is the report. + "fix(#563) paired initials after a comma read as the given name unless a word speaks for them": + _Claim(9, ('_ambiguities', 'family', 'given', 'middle', 'suffix'), "3550ea3b5804", ('DEFAULT',)), # The report-only rule. `_ambiguities` alone, so a widening that # took a ROLE would change the roles here before it reached # the gate -- which is the one thing this row can say about a @@ -6340,8 +6437,11 @@ def _claim(rule: dict) -> _Claim: # `Doe, John DO` moves {family, suffix} here, P6 having # shipped in 2.3; at 2.0.0 and 2.1.0 the same name moves # {middle, suffix} instead. + # 2026-09-30, #563: 15 -> 16, 'García Márquez, Ed G.J.' -- a + # class member in front of a pair keeps the family comma, and + # the pair then ends the given part. Literal-anchored. "fix(#531) a credential ending the given part of a family-comma listing reads as a credential": - _Claim(15, ('_ambiguities', 'middle', 'suffix'), "f17532bb4ff1", ('DEFAULT',)), + _Claim(16, ('_ambiguities', 'middle', 'suffix'), "2da482a17f0e", ('DEFAULT',)), # 2026-09-28, #544: 4 -> 5; gains 'Smith, MD PhD Ma'. "fix(#531) a member the writing declines keeps its name reading and reports the fork": _Claim(5, ('_ambiguities',), "f73bc2fc5408", ('DEFAULT',)), @@ -7894,7 +7994,12 @@ class _Excluded(NamedTuple): #: revisit the day one does. _EXCLUSION_EFFECT: dict[str, _Excluded] = { "(?i)^(?!\\s*ph\\.)(?![^\\s,]+\\s*,\\s*ph\\.\\s*d\\.\\s*$)(?![\\u0000-\\u024f]*\\b(?:jr|sr|ii|iii|iv)\\.?\\s+ph\\.\\s*d\\.\\s*$)[\\u0000-\\u024f]*\\bph\\.\\s*d\\.\\s*$": - _Excluded(6, "69491e3986b1", + _Excluded(7, "85dd96d1f55e", + # 2026-09-30, #563: 6 -> 7 captures, 'John Smith, + # A.B. Ph.D.', a rules.md#C1 boundary line whose + # trailing 'Ph.D.' the end anchor reaches. It reads + # as 1.4.0 read it, so nothing is silenced; + # absorbed_by unchanged. # 2026-09-23, #459: 3 -> 6 captures, rules.md#R4's # three new trailing Ph. D. spellings # ('john smith ph.d.', 'JOHN SMITH PH.D.', diff --git a/tests/v2/test_regex_sync.py b/tests/v2/test_regex_sync.py index fd27f345..4aa99e53 100644 --- a/tests/v2/test_regex_sync.py +++ b/tests/v2/test_regex_sync.py @@ -134,6 +134,9 @@ def test_dotted_initial_is_the_period_alternative_of_initial() -> None: # #544: the chunked sibling of _DOTTED ('M.Eng.'), S2's period gate # for a member with a lower-case tail; no config key to mirror ("_vocab", "_CHUNKED"): None, + # #563: paired initials ('M.J.'), C1's exception to the name-word + # count after a comma; no config key to mirror + ("_vocab", "_PAIRED_INITIALS"): None, ("_group", "_PH"): None, ("_vocab", "_ROMAN"): "roman_numeral", ("_post_rules", "_EAST_SLAVIC"): "east_slavic_patronymic", diff --git a/tools/differential/corpus_rules.jsonl b/tools/differential/corpus_rules.jsonl index d77be1aa..475c5850 100644 --- a/tools/differential/corpus_rules.jsonl +++ b/tools/differential/corpus_rules.jsonl @@ -37,6 +37,8 @@ "Carod i Rovira, Josep" "Carod y de Rovira i" "Davis Royce, Ed" +"De La Cruz, M.J. K.L." +"De La Cruz, M.J. PhD" "Del Toro" "Doe nee Smith Jr. Prof., Jane" "Doe nee Smith Prof. ba" @@ -85,6 +87,11 @@ "Gal·la Serra" "Garcia" "Garcia Juan Carlos" +"García Márquez, Ed G.J." +"García Márquez, G.J" +"García Márquez, G.J." +"García Márquez, G.J.R." +"García Márquez, Ms G.J." "Hans „Erster“ und “Zweiter” Müller" "Hassan Mohamad Ali" "Hassan, Mohamad Ahmad Ali" @@ -193,6 +200,7 @@ "John Smith XYZ" "John Smith Xyz." "John Smith, A.B." +"John Smith, A.B. Ph.D." "John Smith, Ed" "John Smith, Ed Ma" "John Smith, Jones" @@ -208,7 +216,10 @@ "John Smith, Mr. Jr." "John Smith, PhD" "John Smith, PhD MEng" +"John Smith, PhD X.Y." "John Smith, V." +"John Smith, X.Y. P.Q." +"John Smith, X.Y.Z." "John and Jane Smith" "John née Jones Smith MA" "John née Jones Smith Ma" diff --git a/tools/differential/corpus_shapes.jsonl b/tools/differential/corpus_shapes.jsonl index c21d879f..dc9c8aec 100644 --- a/tools/differential/corpus_shapes.jsonl +++ b/tools/differential/corpus_shapes.jsonl @@ -226,7 +226,14 @@ {"name": "doe, jane v phd do", "shape": 2} {"name": "doe, john ma", "shape": 2} {"name": "Davis Royce, Ed", "shape": 3} +{"name": "De La Cruz, M.J. K.L.", "shape": 3} +{"name": "De La Cruz, M.J. PhD", "shape": 3} {"name": "Dr. John P. Doe-Ray, CLU, CFP, LUTC", "shape": 3} +{"name": "García Márquez, Ed G.J.", "shape": 3} +{"name": "García Márquez, G.J.", "shape": 3} +{"name": "García Márquez, G.J.R.", "shape": 3} +{"name": "García Márquez, Ms G.J.", "shape": 3} +{"name": "García Márquez, PhD G.J.", "shape": 3} {"name": "JOHN SMITH, MA", "shape": 3} {"name": "Jane Doe, MS LAc", "shape": 3} {"name": "John Smith Jr., PhD", "shape": 3} @@ -242,6 +249,9 @@ {"name": "John Smith, PhD", "shape": 3} {"name": "John Smith, PhD MEng", "shape": 3} {"name": "John Smith, PhD Ma", "shape": 3} +{"name": "John Smith, PhD X.Y.", "shape": 3} +{"name": "John Smith, X.Y. P.Q.", "shape": 3} +{"name": "John Smith, X.Y.Z.", "shape": 3} {"name": "John Smith, X.Y.Z. MA", "shape": 3} {"name": "Steven Hardman, MD, DO, DDS", "shape": 3} {"name": "john smith, ma", "shape": 3} diff --git a/tools/differential/expected_since_1.4.0.toml b/tools/differential/expected_since_1.4.0.toml index 294c6f19..34268a49 100644 --- a/tools/differential/expected_since_1.4.0.toml +++ b/tools/differential/expected_since_1.4.0.toml @@ -3969,8 +3969,11 @@ issue = "fix(#516) an unlisted dotted acronym is read by position" # reading reported (rules.md#S3). Case says nothing here -- the # periods are the evidence -- which is why 'john smith x.y.z.' moves # with 'John Smith X.Y.Z.'. The same admission reaches the comma form, -# where the NAME-word count decides ('John Smith, A.B.' -> suffix -# 'A.B.'; 'Smith, A.B.' keeps its given and only reports). +# where the NAME-word count decides ('John Smith, X.Y.Z.' -> suffix +# 'X.Y.Z.'; 'Smith, A.B.' keeps its given and only reports). Paired +# initials ('John Smith, A.B.') are the exception to that count and +# have their own rule below (#563); 'García Márquez, G.J.R.' is the +# exception's accepted boundary, three letters reading by the count. # # Two names here restore a reading a PRIOR bundle's vocabulary change # took away, by a different route than it lost it: 'John Smith E.S.Q.' @@ -3999,10 +4002,35 @@ issue = "fix(#516) an unlisted dotted acronym is read by position" # ledgers, for the reason the fix(#289) rule above gives: the class is # a SHAPE, and a regex for the shape would claim every dotted token # the vocabulary already answers for. -name_regex = "^(?:Jack X\\.Y\\.I\\.|John Smith B\\.Tech\\.|John Smith C\\.H\\.A\\.|John Smith E\\.S\\.Q\\.|John Smith Q\\.W\\.E\\.R\\.T\\.|John Smith X\\.Y\\.Z\\.|John Smith, A\\.B\\.|Smith, E\\.S\\.Q\\.|john smith x\\.y\\.z\\.)$" +name_regex = "^(?:García Márquez, G\\.J\\.R\\.|Jack X\\.Y\\.I\\.|John Smith B\\.Tech\\.|John Smith C\\.H\\.A\\.|John Smith E\\.S\\.Q\\.|John Smith Q\\.W\\.E\\.R\\.T\\.|John Smith X\\.Y\\.Z\\.|John Smith, X\\.Y\\.Z\\.|Smith, E\\.S\\.Q\\.|john smith x\\.y\\.z\\.)$" fields = ["family", "given", "middle", "suffix"] orders = ["DEFAULT"] +[[change]] +issue = "fix(#563) paired initials after a comma read as the given name unless a word speaks for them" +# rules.md#C1: "Paired initials are the exception to the count" -- +# two words before the comma may be one surname ('García Márquez, +# G.J.'), so two dotted single letters alone after it stay the given +# name and report the fork, where #516 had read them as a credential. +# "Only a credential in front of them, or another word the class +# admits by its dotted shape standing in the same part, makes them the +# credential run": 'John Smith, PhD X.Y.' and 'John Smith, X.Y. P.Q.' +# move to the run, while 'De La Cruz, M.J. PhD' keeps its initials, +# the degree standing behind them, and so does 'García Márquez, Ms +# G.J.', a title in front ("a suffix word in front of them that is +# not also title vocabulary"), which moves nothing at any baseline. +# The second period is optional ('García Márquez, G.J'). Two pairs +# speaking only for each other make the run and report it ('De La +# Cruz, M.J. K.L.'); a listed member in front speaks for nothing +# ('García Márquez, Ed G.J.' keeps the family comma, and the fix(#531) +# rule claims what its given part's trailing slot then reads). Against a 2.x baseline the names +# whose roles match the baseline's move only the report, which a +# 1.4.0 comparison cannot see; that ledger lists the three role movers +# alone. +name_regex = "^(?:De La Cruz, M\\.J\\. K\\.L\\.|García Márquez, PhD G\\.J\\.|John Smith, PhD X\\.Y\\.|John Smith, X\\.Y\\. P\\.Q\\.)$" +fields = ["family", "given", "middle", "suffix", "title"] +orders = ["DEFAULT"] + [[change]] issue = "fix(given-part-trailing-slot) a credential acronym ending the given part of a family-comma listing reads as a middle name" # 2026-09-19, #531: THE WIDENING THIS RULE'S OWN COMMENT PREDICTED HAS @@ -4078,7 +4106,7 @@ issue = "fix(#531) a dotted credential ending the given part leaves 1.4's middle # dotted token the vocabulary already answers for, which is the reach # fix(#516)'s own comment refuses; _MUST_NOT_MATCH carries the probes # it names. -name_regex = "^Doe, John X\\.Y\\.Z\\.$" +name_regex = "^(?:Doe, John X\\.Y\\.Z\\.|García Márquez, Ed G\\.J\\.)$" fields = ["middle", "suffix"] orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.0.0.toml b/tools/differential/expected_since_2.0.0.toml index 2b5b1f08..c760aa7b 100644 --- a/tools/differential/expected_since_2.0.0.toml +++ b/tools/differential/expected_since_2.0.0.toml @@ -2537,8 +2537,11 @@ issue = "fix(#516) an unlisted dotted acronym is read by position" # reading reported (rules.md#S3). Case says nothing here -- the # periods are the evidence -- which is why 'john smith x.y.z.' moves # with 'John Smith X.Y.Z.'. The same admission reaches the comma form, -# where the NAME-word count decides ('John Smith, A.B.' -> suffix -# 'A.B.'; 'Smith, A.B.' keeps its given and only reports). +# where the NAME-word count decides ('John Smith, X.Y.Z.' -> suffix +# 'X.Y.Z.'; 'Smith, A.B.' keeps its given and only reports). Paired +# initials ('John Smith, A.B.') are the exception to that count and +# have their own rule below (#563); 'García Márquez, G.J.R.' is the +# exception's accepted boundary, three letters reading by the count. # # Two names here restore a reading a PRIOR bundle's vocabulary change # took away, by a different route than it lost it: 'John Smith E.S.Q.' @@ -2567,10 +2570,35 @@ issue = "fix(#516) an unlisted dotted acronym is read by position" # ledgers, for the reason the fix(#289) rule above gives: the class is # a SHAPE, and a regex for the shape would claim every dotted token # the vocabulary already answers for. -name_regex = "^(?:Jack X\\.Y\\.I\\.|John Smith B\\.Tech\\.|John Smith C\\.H\\.A\\.|John Smith E\\.S\\.Q\\.|John Smith Q\\.W\\.E\\.R\\.T\\.|John Smith X\\.Y\\.Z\\.|John Smith, A\\.B\\.|Smith, E\\.S\\.Q\\.|john smith x\\.y\\.z\\.)$" +name_regex = "^(?:García Márquez, G\\.J\\.R\\.|Jack X\\.Y\\.I\\.|John Smith B\\.Tech\\.|John Smith C\\.H\\.A\\.|John Smith E\\.S\\.Q\\.|John Smith Q\\.W\\.E\\.R\\.T\\.|John Smith X\\.Y\\.Z\\.|John Smith, X\\.Y\\.Z\\.|Smith, E\\.S\\.Q\\.|john smith x\\.y\\.z\\.)$" fields = ["family", "given", "middle", "suffix", "_ambiguities"] orders = ["DEFAULT"] +[[change]] +issue = "fix(#563) paired initials after a comma read as the given name unless a word speaks for them" +# rules.md#C1: "Paired initials are the exception to the count" -- +# two words before the comma may be one surname ('García Márquez, +# G.J.'), so two dotted single letters alone after it stay the given +# name and report the fork, where #516 had read them as a credential. +# "Only a credential in front of them, or another word the class +# admits by its dotted shape standing in the same part, makes them the +# credential run": 'John Smith, PhD X.Y.' and 'John Smith, X.Y. P.Q.' +# move to the run, while 'De La Cruz, M.J. PhD' keeps its initials, +# the degree standing behind them, and so does 'García Márquez, Ms +# G.J.', a title in front ("a suffix word in front of them that is +# not also title vocabulary"), which moves nothing at any baseline. +# The second period is optional ('García Márquez, G.J'). Two pairs +# speaking only for each other make the run and report it ('De La +# Cruz, M.J. K.L.'); a listed member in front speaks for nothing +# ('García Márquez, Ed G.J.' keeps the family comma, and the fix(#531) +# rule claims what its given part's trailing slot then reads). Against a 2.x baseline the names +# whose roles match the baseline's move only the report, which a +# 1.4.0 comparison cannot see; that ledger lists the three role movers +# alone. +name_regex = "^(?:De La Cruz, M\\.J\\. K\\.L\\.|De La Cruz, M\\.J\\. PhD|García Márquez, G\\.J|García Márquez, G\\.J\\.|García Márquez, PhD G\\.J\\.|John Smith, A\\.B\\.|John Smith, A\\.B\\. Ph\\.D\\.|John Smith, PhD X\\.Y\\.|John Smith, X\\.Y\\. P\\.Q\\.)$" +fields = ["family", "given", "middle", "suffix", "title", "_ambiguities"] +orders = ["DEFAULT"] + [[change]] issue = "fix(#289/#516) the ambiguous credential class reports at slots that were silent" # Every decision at the trailing suffix slot, the post-comma given @@ -2704,7 +2732,7 @@ issue = "fix(#531) a credential ending the given part of a family-comma listing # measured name by name: what grew is the CORPUS, not the rule, and # a `fix(#533)` rule for them would attribute a #531 reading to the # wrong change. -name_regex = "^(?:DOE, JOHN MA|DOE, MARY JO MA|Doe, Dr\\. John MA|Doe, J\\. MA|Doe, J\\. ba|Doe, John BA|Doe, John MA|Doe, John MA JD|Doe, John MA Jr|Doe, John MA PhD|Doe, John PhD MA|Doe, John Q\\. MA|Doe, John X\\.Y\\.Z\\.|Smith nee Jones, Jane MA|doe, john ma)$" +name_regex = "^(?:DOE, JOHN MA|DOE, MARY JO MA|Doe, Dr\\. John MA|Doe, J\\. MA|Doe, J\\. ba|Doe, John BA|Doe, John MA|Doe, John MA JD|Doe, John MA Jr|Doe, John MA PhD|Doe, John PhD MA|Doe, John Q\\. MA|Doe, John X\\.Y\\.Z\\.|García Márquez, Ed G\\.J\\.|Smith nee Jones, Jane MA|doe, john ma)$" fields = ["_ambiguities", "middle", "suffix"] orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.1.0.toml b/tools/differential/expected_since_2.1.0.toml index 9e8efb93..b4e5381e 100644 --- a/tools/differential/expected_since_2.1.0.toml +++ b/tools/differential/expected_since_2.1.0.toml @@ -2424,8 +2424,11 @@ issue = "fix(#516) an unlisted dotted acronym is read by position" # reading reported (rules.md#S3). Case says nothing here -- the # periods are the evidence -- which is why 'john smith x.y.z.' moves # with 'John Smith X.Y.Z.'. The same admission reaches the comma form, -# where the NAME-word count decides ('John Smith, A.B.' -> suffix -# 'A.B.'; 'Smith, A.B.' keeps its given and only reports). +# where the NAME-word count decides ('John Smith, X.Y.Z.' -> suffix +# 'X.Y.Z.'; 'Smith, A.B.' keeps its given and only reports). Paired +# initials ('John Smith, A.B.') are the exception to that count and +# have their own rule below (#563); 'García Márquez, G.J.R.' is the +# exception's accepted boundary, three letters reading by the count. # # Two names here restore a reading a PRIOR bundle's vocabulary change # took away, by a different route than it lost it: 'John Smith E.S.Q.' @@ -2454,10 +2457,35 @@ issue = "fix(#516) an unlisted dotted acronym is read by position" # ledgers, for the reason the fix(#289) rule above gives: the class is # a SHAPE, and a regex for the shape would claim every dotted token # the vocabulary already answers for. -name_regex = "^(?:Jack X\\.Y\\.I\\.|John Smith B\\.Tech\\.|John Smith C\\.H\\.A\\.|John Smith E\\.S\\.Q\\.|John Smith Q\\.W\\.E\\.R\\.T\\.|John Smith X\\.Y\\.Z\\.|John Smith, A\\.B\\.|Smith, E\\.S\\.Q\\.|john smith x\\.y\\.z\\.)$" +name_regex = "^(?:García Márquez, G\\.J\\.R\\.|Jack X\\.Y\\.I\\.|John Smith B\\.Tech\\.|John Smith C\\.H\\.A\\.|John Smith E\\.S\\.Q\\.|John Smith Q\\.W\\.E\\.R\\.T\\.|John Smith X\\.Y\\.Z\\.|John Smith, X\\.Y\\.Z\\.|Smith, E\\.S\\.Q\\.|john smith x\\.y\\.z\\.)$" fields = ["family", "given", "middle", "suffix", "_ambiguities"] orders = ["DEFAULT"] +[[change]] +issue = "fix(#563) paired initials after a comma read as the given name unless a word speaks for them" +# rules.md#C1: "Paired initials are the exception to the count" -- +# two words before the comma may be one surname ('García Márquez, +# G.J.'), so two dotted single letters alone after it stay the given +# name and report the fork, where #516 had read them as a credential. +# "Only a credential in front of them, or another word the class +# admits by its dotted shape standing in the same part, makes them the +# credential run": 'John Smith, PhD X.Y.' and 'John Smith, X.Y. P.Q.' +# move to the run, while 'De La Cruz, M.J. PhD' keeps its initials, +# the degree standing behind them, and so does 'García Márquez, Ms +# G.J.', a title in front ("a suffix word in front of them that is +# not also title vocabulary"), which moves nothing at any baseline. +# The second period is optional ('García Márquez, G.J'). Two pairs +# speaking only for each other make the run and report it ('De La +# Cruz, M.J. K.L.'); a listed member in front speaks for nothing +# ('García Márquez, Ed G.J.' keeps the family comma, and the fix(#531) +# rule claims what its given part's trailing slot then reads). Against a 2.x baseline the names +# whose roles match the baseline's move only the report, which a +# 1.4.0 comparison cannot see; that ledger lists the three role movers +# alone. +name_regex = "^(?:De La Cruz, M\\.J\\. K\\.L\\.|De La Cruz, M\\.J\\. PhD|García Márquez, G\\.J|García Márquez, G\\.J\\.|García Márquez, PhD G\\.J\\.|John Smith, A\\.B\\.|John Smith, A\\.B\\. Ph\\.D\\.|John Smith, PhD X\\.Y\\.|John Smith, X\\.Y\\. P\\.Q\\.)$" +fields = ["family", "given", "middle", "suffix", "title", "_ambiguities"] +orders = ["DEFAULT"] + [[change]] issue = "fix(#289/#516) the ambiguous credential class reports at slots that were silent" # Every decision at the trailing suffix slot, the post-comma given @@ -2591,7 +2619,7 @@ issue = "fix(#531) a credential ending the given part of a family-comma listing # measured name by name: what grew is the CORPUS, not the rule, and # a `fix(#533)` rule for them would attribute a #531 reading to the # wrong change. -name_regex = "^(?:DOE, JOHN MA|DOE, MARY JO MA|Doe, Dr\\. John MA|Doe, J\\. MA|Doe, J\\. ba|Doe, John BA|Doe, John MA|Doe, John MA JD|Doe, John MA Jr|Doe, John MA PhD|Doe, John PhD MA|Doe, John Q\\. MA|Doe, John X\\.Y\\.Z\\.|Smith nee Jones, Jane MA|doe, john ma)$" +name_regex = "^(?:DOE, JOHN MA|DOE, MARY JO MA|Doe, Dr\\. John MA|Doe, J\\. MA|Doe, J\\. ba|Doe, John BA|Doe, John MA|Doe, John MA JD|Doe, John MA Jr|Doe, John MA PhD|Doe, John PhD MA|Doe, John Q\\. MA|Doe, John X\\.Y\\.Z\\.|García Márquez, Ed G\\.J\\.|Smith nee Jones, Jane MA|doe, john ma)$" fields = ["_ambiguities", "middle", "suffix"] orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.2.0.toml b/tools/differential/expected_since_2.2.0.toml index c5c5d271..bc8254b1 100644 --- a/tools/differential/expected_since_2.2.0.toml +++ b/tools/differential/expected_since_2.2.0.toml @@ -1013,8 +1013,11 @@ issue = "fix(#516) an unlisted dotted acronym is read by position" # reading reported (rules.md#S3). Case says nothing here -- the # periods are the evidence -- which is why 'john smith x.y.z.' moves # with 'John Smith X.Y.Z.'. The same admission reaches the comma form, -# where the NAME-word count decides ('John Smith, A.B.' -> suffix -# 'A.B.'; 'Smith, A.B.' keeps its given and only reports). +# where the NAME-word count decides ('John Smith, X.Y.Z.' -> suffix +# 'X.Y.Z.'; 'Smith, A.B.' keeps its given and only reports). Paired +# initials ('John Smith, A.B.') are the exception to that count and +# have their own rule below (#563); 'García Márquez, G.J.R.' is the +# exception's accepted boundary, three letters reading by the count. # # Two names here restore a reading a PRIOR bundle's vocabulary change # took away, by a different route than it lost it: 'John Smith E.S.Q.' @@ -1043,7 +1046,32 @@ issue = "fix(#516) an unlisted dotted acronym is read by position" # ledgers, for the reason the fix(#289) rule above gives: the class is # a SHAPE, and a regex for the shape would claim every dotted token # the vocabulary already answers for. -name_regex = "^(?:Jack X\\.Y\\.I\\.|John Smith B\\.Tech\\.|John Smith C\\.H\\.A\\.|John Smith E\\.S\\.Q\\.|John Smith Q\\.W\\.E\\.R\\.T\\.|John Smith X\\.Y\\.Z\\.|John Smith, A\\.B\\.|Smith, E\\.S\\.Q\\.|john smith x\\.y\\.z\\.)$" +name_regex = "^(?:García Márquez, G\\.J\\.R\\.|Jack X\\.Y\\.I\\.|John Smith B\\.Tech\\.|John Smith C\\.H\\.A\\.|John Smith E\\.S\\.Q\\.|John Smith Q\\.W\\.E\\.R\\.T\\.|John Smith X\\.Y\\.Z\\.|John Smith, X\\.Y\\.Z\\.|Smith, E\\.S\\.Q\\.|john smith x\\.y\\.z\\.)$" +fields = ["family", "given", "middle", "suffix", "_ambiguities"] +orders = ["DEFAULT"] + +[[change]] +issue = "fix(#563) paired initials after a comma read as the given name unless a word speaks for them" +# rules.md#C1: "Paired initials are the exception to the count" -- +# two words before the comma may be one surname ('García Márquez, +# G.J.'), so two dotted single letters alone after it stay the given +# name and report the fork, where #516 had read them as a credential. +# "Only a credential in front of them, or another word the class +# admits by its dotted shape standing in the same part, makes them the +# credential run": 'John Smith, PhD X.Y.' and 'John Smith, X.Y. P.Q.' +# move to the run, while 'De La Cruz, M.J. PhD' keeps its initials, +# the degree standing behind them, and so does 'García Márquez, Ms +# G.J.', a title in front ("a suffix word in front of them that is +# not also title vocabulary"), which moves nothing at any baseline. +# The second period is optional ('García Márquez, G.J'). Two pairs +# speaking only for each other make the run and report it ('De La +# Cruz, M.J. K.L.'); a listed member in front speaks for nothing +# ('García Márquez, Ed G.J.' keeps the family comma, and the fix(#531) +# rule claims what its given part's trailing slot then reads). Against a 2.x baseline the names +# whose roles match the baseline's move only the report, which a +# 1.4.0 comparison cannot see; that ledger lists the three role movers +# alone. +name_regex = "^(?:De La Cruz, M\\.J\\. K\\.L\\.|De La Cruz, M\\.J\\. PhD|García Márquez, G\\.J|García Márquez, G\\.J\\.|García Márquez, PhD G\\.J\\.|John Smith, A\\.B\\.|John Smith, A\\.B\\. Ph\\.D\\.|John Smith, PhD X\\.Y\\.|John Smith, X\\.Y\\. P\\.Q\\.)$" fields = ["family", "given", "middle", "suffix", "_ambiguities"] orders = ["DEFAULT"] @@ -1172,7 +1200,7 @@ issue = "fix(#531) a credential ending the given part of a family-comma listing # measured name by name: what grew is the CORPUS, not the rule, and # a `fix(#533)` rule for them would attribute a #531 reading to the # wrong change. -name_regex = "^(?:DOE, JOHN MA|DOE, MARY JO MA|Doe, Dr\\. John MA|Doe, J\\. MA|Doe, J\\. ba|Doe, John BA|Doe, John MA|Doe, John MA JD|Doe, John MA Jr|Doe, John MA PhD|Doe, John PhD MA|Doe, John Q\\. MA|Doe, John X\\.Y\\.Z\\.|Smith nee Jones, Jane MA|doe, john ma)$" +name_regex = "^(?:DOE, JOHN MA|DOE, MARY JO MA|Doe, Dr\\. John MA|Doe, J\\. MA|Doe, J\\. ba|Doe, John BA|Doe, John MA|Doe, John MA JD|Doe, John MA Jr|Doe, John MA PhD|Doe, John PhD MA|Doe, John Q\\. MA|Doe, John X\\.Y\\.Z\\.|García Márquez, Ed G\\.J\\.|Smith nee Jones, Jane MA|doe, john ma)$" fields = ["_ambiguities", "middle", "suffix"] orders = ["DEFAULT"] diff --git a/tools/differential/expected_since_2.3.0.toml b/tools/differential/expected_since_2.3.0.toml index b7dc4cea..9f8670fa 100644 --- a/tools/differential/expected_since_2.3.0.toml +++ b/tools/differential/expected_since_2.3.0.toml @@ -301,8 +301,11 @@ issue = "fix(#516) an unlisted dotted acronym is read by position" # reading reported (rules.md#S3). Case says nothing here -- the # periods are the evidence -- which is why 'john smith x.y.z.' moves # with 'John Smith X.Y.Z.'. The same admission reaches the comma form, -# where the NAME-word count decides ('John Smith, A.B.' -> suffix -# 'A.B.'; 'Smith, A.B.' keeps its given and only reports). +# where the NAME-word count decides ('John Smith, X.Y.Z.' -> suffix +# 'X.Y.Z.'; 'Smith, A.B.' keeps its given and only reports). Paired +# initials ('John Smith, A.B.') are the exception to that count and +# have their own rule below (#563); 'García Márquez, G.J.R.' is the +# exception's accepted boundary, three letters reading by the count. # # Two names here restore a reading a PRIOR bundle's vocabulary change # took away, by a different route than it lost it: 'John Smith E.S.Q.' @@ -331,7 +334,32 @@ issue = "fix(#516) an unlisted dotted acronym is read by position" # ledgers, for the reason the fix(#289) rule above gives: the class is # a SHAPE, and a regex for the shape would claim every dotted token # the vocabulary already answers for. -name_regex = "^(?:Jack X\\.Y\\.I\\.|John Smith B\\.Tech\\.|John Smith C\\.H\\.A\\.|John Smith E\\.S\\.Q\\.|John Smith Q\\.W\\.E\\.R\\.T\\.|John Smith X\\.Y\\.Z\\.|John Smith, A\\.B\\.|Smith, E\\.S\\.Q\\.|john smith x\\.y\\.z\\.)$" +name_regex = "^(?:García Márquez, G\\.J\\.R\\.|Jack X\\.Y\\.I\\.|John Smith B\\.Tech\\.|John Smith C\\.H\\.A\\.|John Smith E\\.S\\.Q\\.|John Smith Q\\.W\\.E\\.R\\.T\\.|John Smith X\\.Y\\.Z\\.|John Smith, X\\.Y\\.Z\\.|Smith, E\\.S\\.Q\\.|john smith x\\.y\\.z\\.)$" +fields = ["family", "given", "middle", "suffix", "_ambiguities"] +orders = ["DEFAULT"] + +[[change]] +issue = "fix(#563) paired initials after a comma read as the given name unless a word speaks for them" +# rules.md#C1: "Paired initials are the exception to the count" -- +# two words before the comma may be one surname ('García Márquez, +# G.J.'), so two dotted single letters alone after it stay the given +# name and report the fork, where #516 had read them as a credential. +# "Only a credential in front of them, or another word the class +# admits by its dotted shape standing in the same part, makes them the +# credential run": 'John Smith, PhD X.Y.' and 'John Smith, X.Y. P.Q.' +# move to the run, while 'De La Cruz, M.J. PhD' keeps its initials, +# the degree standing behind them, and so does 'García Márquez, Ms +# G.J.', a title in front ("a suffix word in front of them that is +# not also title vocabulary"), which moves nothing at any baseline. +# The second period is optional ('García Márquez, G.J'). Two pairs +# speaking only for each other make the run and report it ('De La +# Cruz, M.J. K.L.'); a listed member in front speaks for nothing +# ('García Márquez, Ed G.J.' keeps the family comma, and the fix(#531) +# rule claims what its given part's trailing slot then reads). Against a 2.x baseline the names +# whose roles match the baseline's move only the report, which a +# 1.4.0 comparison cannot see; that ledger lists the three role movers +# alone. +name_regex = "^(?:De La Cruz, M\\.J\\. K\\.L\\.|De La Cruz, M\\.J\\. PhD|García Márquez, G\\.J|García Márquez, G\\.J\\.|García Márquez, PhD G\\.J\\.|John Smith, A\\.B\\.|John Smith, A\\.B\\. Ph\\.D\\.|John Smith, PhD X\\.Y\\.|John Smith, X\\.Y\\. P\\.Q\\.)$" fields = ["family", "given", "middle", "suffix", "_ambiguities"] orders = ["DEFAULT"] @@ -460,7 +488,7 @@ issue = "fix(#531) a credential ending the given part of a family-comma listing # measured name by name: what grew is the CORPUS, not the rule, and # a `fix(#533)` rule for them would attribute a #531 reading to the # wrong change. -name_regex = "^(?:DOE, JOHN MA|DOE, MARY JO MA|Doe, Dr\\. John MA|Doe, J\\. MA|Doe, J\\. ba|Doe, John BA|Doe, John MA|Doe, John MA JD|Doe, John MA Jr|Doe, John MA PhD|Doe, John PhD MA|Doe, John Q\\. MA|Doe, John X\\.Y\\.Z\\.|Smith nee Jones, Jane MA|doe, john ma)$" +name_regex = "^(?:DOE, JOHN MA|DOE, MARY JO MA|Doe, Dr\\. John MA|Doe, J\\. MA|Doe, J\\. ba|Doe, John BA|Doe, John MA|Doe, John MA JD|Doe, John MA Jr|Doe, John MA PhD|Doe, John PhD MA|Doe, John Q\\. MA|Doe, John X\\.Y\\.Z\\.|García Márquez, Ed G\\.J\\.|Smith nee Jones, Jane MA|doe, john ma)$" fields = ["_ambiguities", "middle", "suffix"] orders = ["DEFAULT"]