How to read non-english texts in java? They are represented in wrong encoding

2022-12-14 10:02 问答作者：

I use apache HttpClient. And when I'm trying to "read site", all non-english content is represented wrongly.

Actually,开发者_运维知识库 it's represented in windows-1252 but it should be in UTF-8. How can I fix this?

I tried to use InputStreamReader (inputStream, Charset.forName ("UTF-8")), but it didn't help (wrong symbols transformed into ????????).

If the file is in Windows-1252, then telling it to use UTF-8 isn't going to work. Give it Windows-1252 as the charset name, and then you can read the correct data. Knowing what format data should be in isn't nearly as useful as knowing what format it's actually in :)

It's up to you whether you then rewrite it in UTF-8...

Finding the correct character encoding can be a bit of a nightmare. Depending on what the content of your site is, the following might be useful. One thing I've done in the past is rely on a class that will use multiple methods for determining the correct character encoding:

The XmlReader from the rome project will use the UTF byte order mark and/or XML declarations to determine the correct encoding.

So you could use the following construct:

new BufferedReader(new XmlReader(inputStream))

to get to the content.

If the page has encoding in "Content-Type" header, HttpClient will honor it. If not, it will assume Latin-1, not Windows-1252. Are you sure you are getting Windows-1252? You can check encoding like this,

String encoding = method.getResponseCharSet();

If you know the response indeed uses UTF-8 but the header didn't specify it, you can force it to read UTF-8 like this,

byte[] body = method.getResponseBody();
String response = new String(body, "UTF-8");

继续阅读：encoding httpclient internationalization

How to read non-english texts in java? They are represented in wrong encoding

更多精彩内容

精彩评论

最新问答

央视是哪个频道？

请问买过的朋友，舒提啦旅行箱实际使用体验如何？？

检查不孕不育需要的费用？

海信ULED电视画质有什么不同的地方?？

钉子可以挂的住画框幕布吗？

问答排行榜

河神2九牛入海钓河妖是第几集河妖什么来历可活吞牛？

性激素六项检查的最佳时间是多久？多少钱？？

Easiest way to get words of one line from istream into a vector?

《梦在燃烧 (《三国演义》动画片主题曲)》MP3歌词-汤子星？

抽烟只抽炫赫门？