- Package:
- fonts-noto-core
- Source:
- fonts-noto
- Submitter:
- astian
- Date:
- 2020-11-12 23:57:02 UTC
- Severity:
- normal
Dear Maintainer,
With version 20200323-1, when attempting to render code points such as
0x3001 and 0x3002, fontconfig would choose "Noto Sans CJK JP" [0] as
fallback for "Monospace". This was expected behaviour, I want to see
Japanese punctuation glyphs.
0: /usr/share/fonts/opentype/noto/NotoSansCJK-Regular.ttc
Binary packages for version 20200323-1 seem to be gone from the archive
but version 20181227-1 also shows the wanted behaviour.
After updating to version 20201027-3 and later also 20201109-1,
fontconfig chooses "Noto Sans Mongolian" [1]. This results in
unintended glyphs.
1: /usr/share/fonts/truetype/noto/NotoSansMongolian-Regular.ttf
STR:
a) Run:
$ LANG=en_US.UTF-8 pango-view --font monospace -t $'\u3001'
Or even:
$ LANG=en_US.UTF-8 pango-view -t $'\u3001'
b) Run:
$ fc-match --sort monospace family style file | grep -i -e cjk -e mongo
Or even:
$ fc-match --sort : family style file | grep -i -e cjk -e mongo
Expected behaviour:
a) The pango-view window shows the Japanese comma glyph (see for
example "Noto Sans CJK JP" in fontforge).
b) A Japanese font is preferred:
Noto Sans CJK JP:style=Regular:file=/usr/share/fonts/opentype/noto/NotoSansCJK-Regular.ttc
Noto Sans Mongolian:style=Regular:file=/usr/share/fonts/truetype/noto/NotoSansMongolian-Regular.ttf
Actual behaviour:
a) The pango-view window shows a different glyph (from "Noto Sans
Mongolian").
b) A Mongolian font is preferred:
Noto Sans Mongolian:style=Regular:file=/usr/share/fonts/truetype/noto/NotoSansMongolian-Regular.ttf
Noto Sans CJK JP:style=Regular:file=/usr/share/fonts/opentype/noto/NotoSansCJK-Regular.ttc
This looks like a regression and it is one for me, but I guess it could
also be a configuration issue involving fontconfig. I have no custom
fontconfig configuration, though, so if somehow this is not considered a
regression, perhaps you could recommend a configuration that would
restore the previous behaviour for me?
Thanks.
Hi astian, Thanks for a detailed bugreport! Quoting astian (2020-11-11 21:31:00) I understand that this changed. But is it a bug? I mean, is it universally preferred to use Japanese over Mongolian for those characters? one. Failed for me at first, however, so here is another little trick: locale en_US.UTF-8 is not generally enabled, but locale C.UTF-8 is. When I try the above with packages in unstable as of today, I get what looks to me as the comma glyph, even though fc-cache indeed shows Mongolian as prioritized. Sorry, I am not clever with fontconfig and the fonts-noto-core package includes only a small configuration related to older name Droid: /etc/fonts/conf.avail/30-droid-noto.conf I notice that package fonts-noto-cjk ships a more extensive configuration seemingly related to identifying as "monospace": /usr/share/fontconfig/conf.avail/70-fonts-noto-cjk.conf Perhaps it helps to edit that CJK configuration to add binding="strong" also to the monospace sections? Please to report back if that helps, and whether or not you think this is universally a preferred setup or we should perhaps introduce some flexibility in these packages - i.e. a mechanism to let Mongolians prioritize their glyphs and let Japansese prioritize theirs. Kind regards, - Jonas
Jonas Smedegaard:
I forgot about that site. Unfortunately the APT package lists in
stable/testing/unstable no longer have a record of that version and
consequently I no longer have a trusted path to those files.
Universally? How could I say? I know nothing about Mongolian, but
taking U+3001 as example, it's described as "ideographic comma" and
fontconfig (FC) recognises it as part of the charset of various CJK
languages (see for example "fc-lang/{ja,ko,zh_tw,zh_hk}.orth" in FC
sources). Interestingly, FC does not think it is part of mn_cn (see
"fc-lang/mn_cn.orth").
It seems an odd change to me. Here are some cases I can imagine:
- No bug in fonts-noto-core, no technical bug in FC, no policy bug in
FC: I should explicitly configure FC to prioritise a Japanese
fallback font. I should still want an explanation for why is
Mongolian preferred.
- No bug in fonts-noto-core, no technical bug in FC, but policy bug in
FC: Lacking explicit configuration, a Japanese (or some other CJK
language) font should be generally preferred over Mongolian.
- No bug in fonts-noto-core, but technical bug in FC: The changes in
fonts-noto-core have exposed a bug that causes FC to incorrectly
change the fallback selection in favour of this Mongolian font.
- Bug in fonts-noto-core: Either the glyphs don't belong or there is
some incorrect metadata or configuration, or bad interaction with
other fonts, or any combination thereof, which causes FC to
correctly (in technical terms) change the fallback selection in
favour of this Mongolian font.
- Bug in fonts-noto-cjk: As above.
I don't currently have the expertise to decide myself and at present
don't have the time to investigate in depth. Maybe a FC expert's
opinion could help.
Right. I used "en_US" because I wanted an invariable locale, otherwise
the test results would be unreliable because FC determines the default
language from the user's locale and that can affect its decisions. I
briefly looked at the source now and I see that C works too, it also
becomes "en". Thanks.
Here's a Japanese comma (base64, PNG):
iVBORw0KGgoAAAANSUhEUgAAACQAAAAsCAIAAACCMY2iAAAACXBIWXMAAAsTAAALEwEAmpwYAAAA
i0lEQVRYw2P8//8/A70AEwMdwahlo5aNWjZq2ahlo5aNWjZq2ahlo5ZRw7IfNxfEOoRNu/SVHpax
CcrI/NyT5+zVdOQ95e30/wTBv4+nu13FWARtmk98+vefEsBAlKp/H080WvGzaVec/EF7y/7////3
xZmjt75R5LH/jKNdplHLRi0btWzUsoG1DADyRkmF4+MjugAAAABJRU5ErkJggg==
Here's what I get instead (base64, PNG):
iVBORw0KGgoAAAANSUhEUgAAABsAAAAxCAIAAACkgUVSAAAACXBIWXMAAAsTAAALEwEAmpwYAAAA
d0lEQVRIx2P8//8/A1UBEwO1waiJoyaOmjhq4qiJoyaOmjjQJv79+fMPFU3893R9mrFR+sYX//Cp
+k80+Hau1Ypf0DRv/f2f+JQRbeLf2/328q49Zz4TUshIfEvq58cP//gFOAkpYxxtm42aOGriqImj
Jo6aCAcAgcywLcj1CDYAAAAASUVORK5CYII=
Indeed it looks like a comma, but not like a typical Japanese one.
To compare the fonts you can also try fontforge or gucharmap (tick
"View->Show only glyphs from this font").
I tried the suggestion. It did not affect the fallback selection.
I tried blacklisting "Noto Sans Mongolian" and that did work, as can be
expected but it's a very heavy-handed approach. Unfortunately, I
do not yet grok FC configuration to a level that would allow me to
state what I really want:
For code points belonging to the charset of the 'ja' language, always
prioritize a font designed for that language.
Please see above.
Thanks.
astian:
BTW, looking at fontconfig's description of the font (fc-list -v 'Noto Sans
Mongolian') I can see that while it did gain glyphs for new code points, it
already had coverage for U+3001 and U+3002 back in version 20181227-1. Also
"fullname" and "postscriptname" were extended with "Regular". Full diff
against 20201109-1:
Pattern has 24 elts (size 32)
family: "Noto Sans Mongolian"(s)
familylang: "en"(s)
style: "Regular"(s)
stylelang: "en"(s)
- fullname: "Noto Sans Mongolian"(s)
+ fullname: "Noto Sans Mongolian Regular"(s)
fullnamelang: "en"(s)
slant: 0(i)(s)
weight: 80(f)(s)
width: 100(f)(s)
foundry: "GOOG"(s)
file: "/usr/share/fonts/truetype/noto/NotoSansMongolian-Regular.ttf"(s)
index: 0(i)(s)
outline: True(s)
scalable: True(s)
charset:
- 0000: 00000000 80002003 00000000 00000000 00000000 00000001 00000000 00000000
- 0018: 03ff7fff ffffffff ffffffff 00ffffff ffffffff 000007ff 00000000 00000000
- 0020: 00183000 00008000 00000300 00000000 00000000 00000000 00000000 00000000
+ 0000: 00000000 80002307 00000000 00000000 00000000 00000001 00000000 00000000
+ 0018: 03ff7fff ffffffff ffffffff 01ffffff ffffffff 000007ff 00000000 00000000
+ 0020: 30183000 00008000 00000300 00000000 00000000 00000000 00000000 00000000
0024: 00000000 00000000 00000000 000fffff 00000000 00000000 00000000 00000000
0025: 00000000 00000000 00000000 00000000 00000000 00000000 00001000 00000000
- 0030: 00000006 00000000 00000000 00000000 00000000 00000000 00000000 00000000
- 00fe: 00000000 60000000 0000001e 00000000 00000000 00000000 00000000 80000000
+ 0030: 0000fc06 00000000 00000000 00000000 00000000 00000000 00000000 00000000
+ 00fe: 00000000 60000000 0000001e 00000000 00000000 00000000 00000000 00000000
+ 0116: 00000000 00000000 00000000 00001fff 00000000 00000000 00000000 00000000
(s)
lang: mn-cn(s)
- fontversion: 68157(i)(s)
- capability: "otlayout:DFLT otlayout:mong"(s)
+ fontversion: 131138(i)(s)
+ capability: "otlayout:mong"(s)
fontformat: "TrueType"(s)
decorative: False(s)
- postscriptname: "NotoSansMongolian"(s)
+ postscriptname: "NotoSansMongolian-Regular"(s)
color: False(s)
symbol: False(s)
variable: False(s)
Cheers.
Quoting astian (2020-11-12 16:58:00) How you could say? Well, you just did. :-) Thanks for lining up those possibilities. Yeah, gathering some fontconfig experts could help clarify this. Or the more lazy option of simply wait some time, grab another snapshot of Noto, and see if the problem goes away... :-P Thanks. Not convenient for me to extract those images at the moment - may I will later... Let's leave this bugreport open, then. Maybe someone comes by who knows more elegant fontconfig tricks, or one of us gains more knowledge later. Also, I might split the fonts-noto package into smaller parts where it would be possible to simply omit the parts you don't want. - Jonas