#521177 libwww-perl: LWP::UserAgent::request() fails with 'Wide character in syswrite' when posting UTF-8 encoded body

Package:
libwww-perl
Source:
libwww-perl
Submitter:
Ruzsa Balazs
Date:
2026-08-06 21:39:02 UTC
Severity:
important
Tags:
#521177#5
Date:
2009-03-25 13:53:25 UTC
From:
To:
Here is what I tried to do:
------------cut------------
#!/usr/bin/perl

use strict;
use warnings;
use encoding 'iso-8859-2';

use Encode;
use LWP::UserAgent;
use HTTP::Request;

my $POST_URL = "http://somewhere.net/webservice.php";

my $xml = <<"EOT";
<?xml version="1.0" encoding="utf-8" ?>

<PACKET>
<TEXT>Árvíztûrõ tükörfúrógép</TEXT>
</PACKET>
EOT

my $ua = LWP::UserAgent->new();
my $request = HTTP::Request->new('POST', $POST_URL);
my $content = encode('utf-8', $xml);
$request->header('Content-Type' => 'text/xml; charset=utf-8');
$request->header('Content-Length' => length($content));
$request->content($content);
my $response = $ua->request($request);
------------cut------------

Here is what I get when Perl tries to execute the last line:

500 Wide character in syswrite
------------cut------------

The message in the <TEXT> tag is a test phrase containing all possible accented
characters in the Hungarian language. It is encoded as 'iso-8859-2' in the
source file.  Thanks to the 'use encoding' pragma this is converted to
character semantics (utf8 flag on) when Perl reads the source.

After some bughunting, I identified the source of the problem in
/usr/share/perl5/LWP/Protocol/http.pm:

202: my $req_buf = $socket->format_request($method, $fullpath, @h);
...
235: if ($has_content) {
...
249: my $buf = $req_buf . $$content_ref; # <--- HERE

If $$content_ref contains a byte-string (a string with byte semantics) and
$req_buf is a character-string (a string with character semantics) then upon
concatenation, $$content_ref will be converted to character semantics with the
default 'iso-8859-1' encoding (this conversion happens even if $req_buf
contains only ASCII characters). In my example, this means that Perl converts
my utf-8 encoded test phrase to a string that contains consecutive bytes of
utf-8 sequences masquerading as separate characters.

What I don't understand: LWP::UserAgent should be able to send the resulting -
"semantically" wrong, but "syntactically" right - string over the wire, as it
contains only characters with code points < 256. So I still don't understand
where those "wide characters" - which I assume to be characters with code
points >= 256 - are coming from.

Anyway, the problem can be resolved with the following lines added after line
#202:

    my $req_buf = $socket->format_request($method, $fullpath, @h);
    use Encode;
    if (Encode::is_utf8($req_buf)) {
      Encode::_utf8_off($req_buf);
    }

This simply makes sure that the buffer storing the HTTP headers does not have
the 'utf8' flag turned on. I can only hope that the $req_buf returned by
format_request does not contain non-ASCII characters (it shouldn't).

With this change, the concatenation above does not touch $$content_ref and the
request gets posted without errors.

#521177#10
Date:
2009-03-26 08:58:40 UTC
From:
To:
Net::HTTP::Methods line #167:

...
return join($CRLF, "$method $uri HTTP/$ver", @h2, @h, "", $content);

If $content is a string of bytes (as opposed to a string of characters),
then joining them with the rest (which are strings of characters) will
do an Encode::decode('iso-8859-1', ...) on $content and use the result
in the concatenation.

#521177#25
Date:
2026-08-06 21:36:51 UTC
From:
To:
Control: tag -1 + wontfix
https://github.com/libwww-perl/libwww-perl/issues/205

So closing in Debian as well.

Cheers,
gregor