Skip to content

Preserve lowercase particles in two-word names - #627

Open
natesute wants to merge 2 commits into
sciunto-org:mainfrom
natesute:fix/two-word-name-particles
Open

natesute wants to merge 2 commits into
sciunto-org:mainfrom
natesute:fix/two-word-name-particles

Conversation

@natesute

Copy link
Copy Markdown

Fixes #626.

The two-word shortcut assumes the first token is a first name, so rewriting de Gaulle and van Gogh in last-name-first style produces Gaulle, de and Gogh, van. This changes how BibTeX abbreviates and sorts those names.

Remove the shortcut so two-word names use the existing case-aware partitioning. The shared parser/inverse test cases now cover lowercase particles, two lowercase words, lowercase and uppercase LaTeX accents, and a brace-protected token. A parse/write/parse test checks the full middleware round trip.

Validation: 21 regression cases fail on the base. The full test suite passes 2,806 tests with 12 skips. All pre-commit hooks pass on the changed files. The requested all-files run also passed the Python/style hooks, but the TOML formatter wanted unrelated whitespace changes in the existing pyproject.toml; those were left out.

@MiWeiss

MiWeiss commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

Thanks! The fix for de Gaulle / van Gogh is correct and matches BibTeX.

One problem: the two-word shortcut was also hiding gaps in how we detect a word's case. Removing it exposes them for the most common name shape. These now parse as von + last instead of first + last:

parse_single_name_into_parts(r"{\L}ukasz Chmielewski")
parse_single_name_into_parts(r"{\O}yvind Ytrehus")
parse_single_name_into_parts(r"{\'{E}}ric Brier")           # nested accent, dblp style
parse_single_name_into_parts(r"{Jean-Fran\c{c}ois} Biasse")

BibTeX treats all four as first + last, for three reasons:

  • It counts \L, \O, \AA, \AE and \OE as uppercase.
  • It reads the letter inside nested braces like {\'{E}}.
  • It never treats a brace group that doesn't start with \ as von.

On cryptobib, 47 names change: 31 regress like the examples above, and 16 are lowercase given names that now follow BibTeX. Performance is unchanged.

Could you fix the case detection for these shapes in this PR and add them to the test cases?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SplitNameParts: two-word names like "van Gogh" or "de Gaulle" put the particle in first instead of von

2 participants