Repository navigation
Conversation
Collaborator
|
Thanks! The fix for One problem: the two-word shortcut was also hiding gaps in how we detect a word's case. Removing it exposes them for the most common name shape. These now parse as parse_single_name_into_parts(r"{\L}ukasz Chmielewski")
parse_single_name_into_parts(r"{\O}yvind Ytrehus")
parse_single_name_into_parts(r"{\'{E}}ric Brier") # nested accent, dblp style
parse_single_name_into_parts(r"{Jean-Fran\c{c}ois} Biasse")BibTeX treats all four as first + last, for three reasons:
On cryptobib, 47 names change: 31 regress like the examples above, and 16 are lowercase given names that now follow BibTeX. Performance is unchanged. Could you fix the case detection for these shapes in this PR and add them to the test cases? |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #626.
The two-word shortcut assumes the first token is a first name, so rewriting
de Gaulle and van Goghin last-name-first style producesGaulle, de and Gogh, van. This changes how BibTeX abbreviates and sorts those names.Remove the shortcut so two-word names use the existing case-aware partitioning. The shared parser/inverse test cases now cover lowercase particles, two lowercase words, lowercase and uppercase LaTeX accents, and a brace-protected token. A parse/write/parse test checks the full middleware round trip.
Validation: 21 regression cases fail on the base. The full test suite passes 2,806 tests with 12 skips. All pre-commit hooks pass on the changed files. The requested all-files run also passed the Python/style hooks, but the TOML formatter wanted unrelated whitespace changes in the existing pyproject.toml; those were left out.