镜像站点 · 本页由第三方 GitHub 只读镜像提供,非 GitHub 官方站点,不接受任何登录或凭据输入。前往 github.com
Skip to content

Docs: subprocess text mode does not say that the default encoding is a guess about the child process #158633

Description

@AlKor13

Documentation

subprocess's text-mode documentation says which encoding is used, but not that the value is a guess about the child process. In gh-105312 @zooba said a docs enhancement would be considered, and described what it should say:

I would add a warning that it will guess the encoding of the child process and you'll get garbage if it guesses wrong, and you are always recommended to specify the encoding to use (on all platforms, though you are most likely to get into trouble on Windows).

and, on the Windows specifics:

It's worse on Windows, because there have been 2-3 common defaults over time, and even GetConsoleOutputCP is only correct if you know that the child process uses it.

This is the follow-up issue for that. @kunom, who filed gh-105312, supported it there.

What the docs say today

Doc/library/subprocess.rst, the frequently-used-arguments section:

If encoding or errors are specified, or text (also known as universal_newlines) is true, the file objects stdin, stdout and stderr will be opened in text mode using the encoding and errors specified in the call or the defaults for :class:io.TextIOWrapper.

Accurate, and it never says that the default is unrelated to what the child writes, or what the failure looks like when they disagree.

Why the failure is hard to attribute

Measured on Windows 11, Python 3.12.8, ANSI code page cp1251, console output code page cp866 — two different defaults in force at the same time, so no single value is correct for every child of one process:

subprocess.run([sys.executable, "-c", "print('Тест')"], capture_output=True, text=True).stdout
# 'Тест'                      the child wrote UTF-8, text=True decoded cp1251

subprocess.run("echo Тест", shell=True, text=True, stdout=subprocess.PIPE).stdout
# '’Ґбв'                          the child wrote cp866, text=True decoded cp1251

Under UTF-8 mode, which :pep:686 makes the default in 3.15, the second one stops being wrong text and becomes an exception instead:

PYTHONUTF8=1 -> UnicodeDecodeError: 'utf-8' codec can't decode byte 0x92 in position 0

The traceback for that comes from inside the stream reader rather than from the caller's code, which is what sends people looking for a bug in subprocess rather than at their own missing encoding=. That is the part a warning in the docs would save.

Proposed change

A .. warning:: in frequently-used-arguments, right after the paragraph that describes binary mode, keeping to @zooba's framing: the default is a guess, specify encoding on every platform, and on Windows a console child is read with the console output code page while a Python child can be told what to use through :envvar:PYTHONUTF8 / :envvar:PYTHONIOENCODING. No behaviour change, no new API, and Popen's section already points at this one.

Happy to open the PR — a draft is ready and I will link it here.

Linked PRs

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    docsDocumentation in the Doc dirtopic-subprocessSubprocess issues.

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions