Repository navigation
Recognize multi-word credential suffixes after a comma ("Smith, LEED AP") #291
Description
Activity
Residual folded in from #274: the Polish maiden marker "z domu" (lit. "of the house [of]") was descoped from 2.0 because it is a two-token marker and the vocabulary matches one word at a time — the same limitation this issue covers for credential suffixes (
LEED AP). The descope is currently noted only inconfig/maiden_markers.py's docstring.Whatever design lands here (segment-level matching on the comma path, or whitespace-permitted vocabulary entries with phrase matching), it should decide whether multi-token marker phrases ride the same mechanism — "Maria Kowalska z domu Nowak" is the test case.
- added a commit that references this issue
on Aug 22, 2026 Closing as working-as-designed. The premise this issue rests on — and that the approved bundle spec built decision 2 on — does not hold.
Space-separated post-nominal runs already parse
Stock master, no changes:
parse("John Smith, MD PhD").suffix # 'MD PhD' parse("John Smith, CBE MC").suffix # 'CBE MC' parse("John Smith, BSc MBA").suffix # 'BSc MBA' parse("John Smith, USN Ret.").suffix # 'USN Ret.' parse("John Smith, PSM I").suffix # 'PSM I'
And this is not a recent fix — measured on the wheels:
input 1.4.0 2.0.0 2.1.0 master John Smith, MD PhD'MD PhD''MD PhD''MD PhD''MD PhD'John Smith, PSM I'PSM I''PSM I''PSM I''PSM I'It is true that vocabulary is matched one token at a time, so a stored
"leed ap"is inert. What does not follow is that the shape is unreachable: the run predicateis_wholly_suffix(nameparser/_pipeline/_vocab.py:246) reassembles adjacent suffix tokens after matching, so a multi-word credential is reachable as its component words.Two of the seven "dead" entries prove it.
psmis already inSUFFIX_ACRONYMSandi/iiare already inSUFFIX_WORDS, sopsm iandpsm iiparse today with no changes at all.LEED APfails only becauseleedandapare absent from the vocabulary:from nameparser import Lexicon, Parser parser = Parser(lexicon=Lexicon.default().add(suffix_acronyms={"leed", "ap"})) parser.parse("John Smith, LEED AP").suffix # 'LEED AP'
Why those words are not being added to the shipped vocabulary
leedis borne as a surname, and the set it would join has an unaudited collision surface: 575 of 579 alphabeticSUFFIX_ACRONYMSentries leavefamilyempty in"John <word>", against exactly 4 ambiguous-gated exceptions (ma,do,ed,jd). Un-gated entries that are real surnames today includerai(#342, corpus-attested viaAishwarya Rai),ba,cha,sa,se,omandmc.Shipping
leedandapwould add to that surface to serve a credential with no attested demand — the onlyLEED APstrings in any corpus are the ones this issue introduced, and the same is true of theJohn LeedandMary Nicetcounter-examples. A caller who genuinely parses LEED credentials adds two words to aLexiconand gets the existing machinery.What falls away with this
SUFFIX_PHRASES, the new segment-level matching unit, theis_suffix_phrasepredicate, and theLexicon.suffix_phrasesfield — commit 4 of the approved bundle plan.- Amendment A6 (the glued honorific peel stepping over a phrase segment). With no phrase-matching unit,
_is_post_nominal's token-level test stays correct by construction. The김민준씨, LEED APexample that drove it does not survive scrutiny independently: CLDR'ssortingpatterns forkoandjacarry no comma at all (a surname-first locale is already in sorting order, so there is nothing to invert), and씨is specifically the honorific for someone without a professional title. - The multi-word
UserWarningneeds no carve-out and stays correct as written.
The rest of the bundle shipped in #428 (#296 / #325) and stands.
Spun off
- customize.rst says a multi-word entry can never match, but never says the credential is still reachable #433 — the docs assert this limitation and should document the run behavior instead.
- parse("Maria Kowalska z domu Nowak") cannot reach the Polish maiden marker — markers do not compose the way suffixes do #434 — "z domu", folded in here on 2026-07-27, does not dissolve the same way. Measured: adding
zanddomutomaiden_markersgivesmaiden='domu Nowak', because markers have no run predicate. It needs its own decision. Raiis parsed as a post-nominal suffix, consuming a common South Asian surname #342 will carry thedocs/design/decisions.mdamendment to the comma-suffix-arc Declined entry. The split-into-single-words decline was correct as measured; what was wrong was inferring an unreachable shape from an inert entry.
- added a commit that references this issue
on Aug 24, 2026 - added 4 commits that reference this issue
on Aug 29, 2026 - added a commit that references this issue
on Sep 1, 2026 - added a commit that references this issue
on Sep 22, 2026 - added 4 commits that reference this issue
on Oct 8, 2026
"John Smith, LEED AP"parses with given=LEED— the credential is read as a name. The vocabulary entries that were meant to cover this (leed ap,nicet i–nicet iv,psm i,psm ii) could never match in any release: vocabulary is matched one word at a time, so a multi-word entry is inert (verified empirically against 1.4.0 on PyPI, and those dead entries were removed in 2.0.0).Splitting them into single-word entries is disqualified by collisions, verified during the 2.0 API review:
Smith, A.P.A.P.A.P.John LeedLeed, family lostLeedMary NicetNicet, family lostNicetThe period-gated
suffix_acronyms_ambiguousescape doesn't help either — nobody writesL.E.E.D., so gating on periods is equivalent to removal.The workable shape: on the suffix-comma path, match the whole comma segment against suffix vocabulary as a unit. The segment arrives as one piece there (
"Smith, LEED AP"→ segmentLEED AP), so multi-word credentials become recognizable without touching no-comma parsing or reintroducing the collision surface. This would need a decision about which vocabulary field holds multi-word credentials (a new segment-matched set, or allowing whitespace insuffix_acronymswith segment-level matching) — the current multi-wordUserWarningwould need to carve out whichever home is chosen.