Documentation
subprocess's text-mode documentation says which encoding is used, but not that the value is a guess about the child process. In gh-105312 @zooba said a docs enhancement would be considered, and described what it should say:
I would add a warning that it will guess the encoding of the child process and you'll get garbage if it guesses wrong, and you are always recommended to specify the encoding to use (on all platforms, though you are most likely to get into trouble on Windows).
and, on the Windows specifics:
It's worse on Windows, because there have been 2-3 common defaults over time, and even GetConsoleOutputCP is only correct if you know that the child process uses it.
This is the follow-up issue for that. @kunom, who filed gh-105312, supported it there.
What the docs say today
Doc/library/subprocess.rst, the frequently-used-arguments section:
If encoding or errors are specified, or text (also known as universal_newlines) is true, the file objects stdin, stdout and stderr will be opened in text mode using the encoding and errors specified in the call or the defaults for :class:io.TextIOWrapper.
Accurate, and it never says that the default is unrelated to what the child writes, or what the failure looks like when they disagree.
Why the failure is hard to attribute
Measured on Windows 11, Python 3.12.8, ANSI code page cp1251, console output code page cp866 — two different defaults in force at the same time, so no single value is correct for every child of one process:
subprocess.run([sys.executable, "-c", "print('Тест')"], capture_output=True, text=True).stdout
# 'Тест' the child wrote UTF-8, text=True decoded cp1251
subprocess.run("echo Тест", shell=True, text=True, stdout=subprocess.PIPE).stdout
# '’Ґбв' the child wrote cp866, text=True decoded cp1251
Under UTF-8 mode, which :pep:686 makes the default in 3.15, the second one stops being wrong text and becomes an exception instead:
PYTHONUTF8=1 -> UnicodeDecodeError: 'utf-8' codec can't decode byte 0x92 in position 0
The traceback for that comes from inside the stream reader rather than from the caller's code, which is what sends people looking for a bug in subprocess rather than at their own missing encoding=. That is the part a warning in the docs would save.
Proposed change
A .. warning:: in frequently-used-arguments, right after the paragraph that describes binary mode, keeping to @zooba's framing: the default is a guess, specify encoding on every platform, and on Windows a console child is read with the console output code page while a Python child can be told what to use through :envvar:PYTHONUTF8 / :envvar:PYTHONIOENCODING. No behaviour change, no new API, and Popen's section already points at this one.
Happy to open the PR — a draft is ready and I will link it here.
Linked PRs
Documentation
subprocess's text-mode documentation says which encoding is used, but not that the value is a guess about the child process. In gh-105312 @zooba said a docs enhancement would be considered, and described what it should say:and, on the Windows specifics:
This is the follow-up issue for that. @kunom, who filed gh-105312, supported it there.
What the docs say today
Doc/library/subprocess.rst, the frequently-used-arguments section:Accurate, and it never says that the default is unrelated to what the child writes, or what the failure looks like when they disagree.
Why the failure is hard to attribute
Measured on Windows 11, Python 3.12.8, ANSI code page
cp1251, console output code pagecp866— two different defaults in force at the same time, so no single value is correct for every child of one process:Under UTF-8 mode, which :pep:
686makes the default in 3.15, the second one stops being wrong text and becomes an exception instead:The traceback for that comes from inside the stream reader rather than from the caller's code, which is what sends people looking for a bug in
subprocessrather than at their own missingencoding=. That is the part a warning in the docs would save.Proposed change
A
.. warning::in frequently-used-arguments, right after the paragraph that describes binary mode, keeping to @zooba's framing: the default is a guess, specify encoding on every platform, and on Windows a console child is read with the console output code page while a Python child can be told what to use through :envvar:PYTHONUTF8/ :envvar:PYTHONIOENCODING. No behaviour change, no new API, andPopen's section already points at this one.Happy to open the PR — a draft is ready and I will link it here.
Linked PRs