#988554 xterm does not render U+A7BA ... U+A7F6

Package:
xterm
Source:
xterm
Description:
X terminal emulator
Submitter:
Uwe Waldmann
Date:
2022-02-14 00:00:05 UTC
Severity:
normal
Tags:
#988554#5
Date:
2021-05-15 14:11:12 UTC
From:
To:
Dear Maintainer,

xterm renders all characters between U+A7BA and U+A7F6 with width zero,
that is, not at all, even if they are present in the chosen font.
For instance,

  printf '<\uA7B9>\n'

displays "<u>" with a slash through "u" (correct), whereas

  printf '<\uA7BA>\n'

displays "<>" (incorrect) instead of "<A>" with an apostrophe before
"A". The problem shows up with arbitrary fonts (both pcf and otb are
affected). Using -mk_width does not change the situation.

(In Debian 9.1, it still worked correctly.)

Expected behavior:

Characters between U+A7BA and U+A7F6 should be rendered properly
if present in the given font(s), or replaced by a default character
otherwise.

#988554#10
Date:
2021-05-16 23:29:25 UTC
From:
To:
According to https://unicodeplus.com/U+A7BA

	The character Ꞻ (Latin Capital Letter Glottal A) is represented by the
	Unicode codepoint U+A7BA.  It is encoded in the Latin Extended-D block,
	which belongs to the Basic Multilingual Plane.  It was added to Unicode
	in version 12.0 (March, 2019).  It is HTML encoded as Ꞻ.

xterm #344 is a little earlier than that.  Its fallback copy of wcwidth
doesn't list that range (I updated the table to Unicode 12 in #345,
and added a test-driver around that time).

The system wcwidth doesn't cover that range either.

Characters which aren't known to wcwidth are treated as nonprinting...

hmm - which version of xterm was that?

I'm guessing that it was #327

(it should not have worked, but there's always the possibility that I
fixed a bug which was making it appear to work)

#988554#13
Date:
2021-05-16 23:29:25 UTC
From:
To:
According to https://unicodeplus.com/U+A7BA

	The character Ꞻ (Latin Capital Letter Glottal A) is represented by the
	Unicode codepoint U+A7BA.  It is encoded in the Latin Extended-D block,
	which belongs to the Basic Multilingual Plane.  It was added to Unicode
	in version 12.0 (March, 2019).  It is HTML encoded as Ꞻ.

xterm #344 is a little earlier than that.  Its fallback copy of wcwidth
doesn't list that range (I updated the table to Unicode 12 in #345,
and added a test-driver around that time).

The system wcwidth doesn't cover that range either.

Characters which aren't known to wcwidth are treated as nonprinting...

hmm - which version of xterm was that?

I'm guessing that it was #327

(it should not have worked, but there's always the possibility that I
fixed a bug which was making it appear to work)

#988554#18
Date:
2021-05-17 07:22:35 UTC
From:
To:
yes.

OK, that's possible. Thanks for the explanation.

Best

Uwe

#988554#23
Date:
2021-05-17 22:58:20 UTC
From:
To:
In #327, xterm's wcwidth checked if the codes were combining characters
(using a table), or control characters and (for example this case) matched it
against some ranges of double-width characters.  If it was none of those, it
assumed single-width.

Starting in #330, I added another table "unknowns" to account for
codes which had no specific width:

Patch #330 - 2017/06/20
     * modify wcwidth.c to return -1 for non-Unicode values, and adjust a
       couple of blocks to better match assumptions about ambiguous-width
       characters  in  other  implementations.  Also  modify wcwidth.c to
       support configurable soft-hyphen, so there is no drawback to using
       this version rather than a system wcwidth.

#988554#26
Date:
2021-05-17 22:58:20 UTC
From:
To:
In #327, xterm's wcwidth checked if the codes were combining characters
(using a table), or control characters and (for example this case) matched it
against some ranges of double-width characters.  If it was none of those, it
assumed single-width.

Starting in #330, I added another table "unknowns" to account for
codes which had no specific width:

Patch #330 - 2017/06/20
     * modify wcwidth.c to return -1 for non-Unicode values, and adjust a
       couple of blocks to better match assumptions about ambiguous-width
       characters  in  other  implementations.  Also  modify wcwidth.c to
       support configurable soft-hyphen, so there is no drawback to using
       this version rather than a system wcwidth.