镜像站点 · 本页由第三方 GitHub 只读镜像提供,非 GitHub 官方站点,不接受任何登录或凭据输入。前往 github.com
Skip to content

Recognize multi-word credential suffixes after a comma ("Smith, LEED AP") #291

Description

@derek73

"John Smith, LEED AP" parses with given=LEED — the credential is read as a name. The vocabulary entries that were meant to cover this (leed ap, nicet i–nicet iv, psm i, psm ii) could never match in any release: vocabulary is matched one word at a time, so a multi-word entry is inert (verified empirically against 1.4.0 on PyPI, and those dead entries were removed in 2.0.0).

Splitting them into single-word entries is disqualified by collisions, verified during the 2.0 API review:

Input With split entries Today (correct)
Smith, A.P. suffix A.P. given A.P.
John Leed suffix Leed, family lost family Leed
Mary Nicet suffix Nicet, family lost family Nicet

The period-gated suffix_acronyms_ambiguous escape doesn't help either — nobody writes L.E.E.D., so gating on periods is equivalent to removal.

The workable shape: on the suffix-comma path, match the whole comma segment against suffix vocabulary as a unit. The segment arrives as one piece there ("Smith, LEED AP" → segment LEED AP), so multi-word credentials become recognizable without touching no-comma parsing or reintroducing the collision surface. This would need a decision about which vocabulary field holds multi-word credentials (a new segment-matched set, or allowing whitespace in suffix_acronyms with segment-level matching) — the current multi-word UserWarning would need to carve out whichever home is chosen.

Activity

  1. added this to the v2.1 milestone on Jul 28, 2026
  2. derek73 commented on Jul 28, 2026

    @derek73
    OwnerAuthor

    Residual folded in from #274: the Polish maiden marker "z domu" (lit. "of the house [of]") was descoped from 2.0 because it is a two-token marker and the vocabulary matches one word at a time — the same limitation this issue covers for credential suffixes (LEED AP). The descope is currently noted only in config/maiden_markers.py's docstring.

    Whatever design lands here (segment-level matching on the comma path, or whitespace-permitted vocabulary entries with phrase matching), it should decide whether multi-token marker phrases ride the same mechanism — "Maria Kowalska z domu Nowak" is the test case.

  3. modified the milestones: v2.1, v2.2 on Aug 1, 2026
  4. self-assigned this
    on Aug 23, 2026
  5. derek73 commented on Aug 24, 2026

    @derek73
    OwnerAuthor

    Closing as working-as-designed. The premise this issue rests on — and that the approved bundle spec built decision 2 on — does not hold.

    Space-separated post-nominal runs already parse

    Stock master, no changes:

    parse("John Smith, MD PhD").suffix    # 'MD PhD'
    parse("John Smith, CBE MC").suffix    # 'CBE MC'
    parse("John Smith, BSc MBA").suffix   # 'BSc MBA'
    parse("John Smith, USN Ret.").suffix  # 'USN Ret.'
    parse("John Smith, PSM I").suffix     # 'PSM I'

    And this is not a recent fix — measured on the wheels:

    input 1.4.0 2.0.0 2.1.0 master
    John Smith, MD PhD 'MD PhD' 'MD PhD' 'MD PhD' 'MD PhD'
    John Smith, PSM I 'PSM I' 'PSM I' 'PSM I' 'PSM I'

    It is true that vocabulary is matched one token at a time, so a stored "leed ap" is inert. What does not follow is that the shape is unreachable: the run predicate is_wholly_suffix (nameparser/_pipeline/_vocab.py:246) reassembles adjacent suffix tokens after matching, so a multi-word credential is reachable as its component words.

    Two of the seven "dead" entries prove it. psm is already in SUFFIX_ACRONYMS and i/ii are already in SUFFIX_WORDS, so psm i and psm ii parse today with no changes at all.

    LEED AP fails only because leed and ap are absent from the vocabulary:

    from nameparser import Lexicon, Parser
    parser = Parser(lexicon=Lexicon.default().add(suffix_acronyms={"leed", "ap"}))
    parser.parse("John Smith, LEED AP").suffix   # 'LEED AP'

    Why those words are not being added to the shipped vocabulary

    leed is borne as a surname, and the set it would join has an unaudited collision surface: 575 of 579 alphabetic SUFFIX_ACRONYMS entries leave family empty in "John <word>", against exactly 4 ambiguous-gated exceptions (ma, do, ed, jd). Un-gated entries that are real surnames today include rai (#342, corpus-attested via Aishwarya Rai), ba, cha, sa, se, om and mc.

    Shipping leed and ap would add to that surface to serve a credential with no attested demand — the only LEED AP strings in any corpus are the ones this issue introduced, and the same is true of the John Leed and Mary Nicet counter-examples. A caller who genuinely parses LEED credentials adds two words to a Lexicon and gets the existing machinery.

    What falls away with this

    • SUFFIX_PHRASES, the new segment-level matching unit, the is_suffix_phrase predicate, and the Lexicon.suffix_phrases field — commit 4 of the approved bundle plan.
    • Amendment A6 (the glued honorific peel stepping over a phrase segment). With no phrase-matching unit, _is_post_nominal's token-level test stays correct by construction. The 김민준씨, LEED AP example that drove it does not survive scrutiny independently: CLDR's sorting patterns for ko and ja carry no comma at all (a surname-first locale is already in sorting order, so there is nothing to invert), and 씨 is specifically the honorific for someone without a professional title.
    • The multi-word UserWarning needs no carve-out and stays correct as written.

    The rest of the bundle shipped in #428 (#296 / #325) and stands.

    Spun off

    Also filed while measuring this: #429, #430, #431, #432.

  6. added a commit that references this issue on Sep 22, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Projects

No projects

    Milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions