当初在设计weblucene 的时候,为了能够正确的截取请求中的中文q参数,在执行request.getparameter("q")之前先调用了request.setcharacterencoding("gb2312")方法。这样虽然避免了乱码问题的出现,但却使得weblucene 同时只能对一种编码进行处理,无法实现类似于google 的搜索效果,例如下面两个链接:
http://www.google.com/search?ie=utf-8&q=%e5%8c%97%e4%ba%ac
http://www.google.com/search?ie=gb2312&q=%b1%b1%be%a9
http://www.google.com/search?ie=utf-8&q=%e5%8c%97%e4%ba%ac
http://www.google.com/search?ie=gb2312&q=%b1%b1%be%a9
查找的都是关键字为“北京”的内容,只是前者的 url 是以utf-8 方式进行urlencode 得的,而后者则采用的是gb2312。
在google 上查找了很多关于setcharacterencoding 和“多语言支持” 的文章,最后还是用最古老的方法解决了问题,见下面一段代码:
string ie = request.getparameter("encoding");
string q = new string(request.getparameter("q").getbytes("iso-8859-1"), ie);
这样,只要$q 是以$encodig 方式进行提交/urlencode 的,那么我们就可以得到正确的q参数了——但有一点需要注意,就是要用 “iso-8859-1” 的形式对字符串q 进行getbytes() 操作,而不能采用utf-8的形式(至少我目前的试验结果是这样的)。
之所以不能用utf-8 是因为一些应用服务器,如resin2.*,在试图用utf-8 编码解析参数的时候,往往会因为参数中含有了不合法的utf-8 字符而产生异常,从而导致无法正常的解析参数。
这样,只要$q 是以$encodig 方式进行提交/urlencode 的,那么我们就可以得到正确的q参数了——但有一点需要注意,就是要用 “iso-8859-1” 的形式对字符串q 进行getbytes() 操作,而不能采用utf-8的形式(至少我目前的试验结果是这样的)。
之所以不能用utf-8 是因为一些应用服务器,如resin2.*,在试图用utf-8 编码解析参数的时候,往往会因为参数中含有了不合法的utf-8 字符而产生异常,从而导致无法正常的解析参数。
稍后请见详细描述。
大家可以用我后面的程序作一个试验:
首先把resin的缺省编码设为utf-8:
<log id='/' href='stderr:' timestamp="[%h:%m:%s.%s]"/>
<web-app character-encoding='utf-8'>
....
<servlet>
<servlet-name>character</servlet-name>
<servlet-class>com.chedong.weblucene.characterencodingtest</servlet-class>
</servlet>
<servlet-mapping>
<servlet-name>character</servlet-name>
<url-pattern>/character</url-pattern>
</servlet-mapping>
....
</web-app>
然后访问http://localhost:8080/weblucene/search?encoding=utf-8&q=%e5%8c%97%e4%ba%ac,在resin 的log 中你会发现:
[19:10:52.484] java.io.charconversionexception: illegal utf8 encoding
at com.caucho.vfs.i18n.utf8reader.read(utf8reader.java:102)
at com.caucho.vfs.bytetochar.readchar(bytetochar.java:179)
at com.caucho.vfs.bytetochar.getconvertedstring(bytetochar.java:126)
at com.caucho.server.http.form.parsequerystring(form.java:100)
at com.caucho.server.http.request.parsequery(request.java:1352)
at com.caucho.server.http.request.getparametervalues(request.java:1449)
at com.caucho.server.http.request.getparameter(request.java:1459)
at com.chedong.weblucene.characterencodingtest.doget(unknown source)
at javax.servlet.http.httpservlet.service(httpservlet.java:126)
at javax.servlet.http.httpservlet.service(httpservlet.java:103)
at com.caucho.server.http.filterchainservlet.dofilter(filterchainservlet.java:96)
at com.caucho.server.http.invocation.service(invocation.java:315)
at com.caucho.server.http.cacheinvocation.service(cacheinvocation.java:135)
at com.caucho.server.http.httprequest.handlerequest(httprequest.java:246)
at com.caucho.server.http.httprequest.handleconnection(httprequest.java:163)
at com.caucho.server.tcpconnection.run(tcpconnection.java:139)
at java.lang.thread.run(thread.java:536)
<log id='/' href='stderr:' timestamp="[%h:%m:%s.%s]"/>
<web-app character-encoding='utf-8'>
....
<servlet>
<servlet-name>character</servlet-name>
<servlet-class>com.chedong.weblucene.characterencodingtest</servlet-class>
</servlet>
<servlet-mapping>
<servlet-name>character</servlet-name>
<url-pattern>/character</url-pattern>
</servlet-mapping>
....
</web-app>
然后访问http://localhost:8080/weblucene/search?encoding=utf-8&q=%e5%8c%97%e4%ba%ac,在resin 的log 中你会发现:
[19:10:52.484] java.io.charconversionexception: illegal utf8 encoding
at com.caucho.vfs.i18n.utf8reader.read(utf8reader.java:102)
at com.caucho.vfs.bytetochar.readchar(bytetochar.java:179)
at com.caucho.vfs.bytetochar.getconvertedstring(bytetochar.java:126)
at com.caucho.server.http.form.parsequerystring(form.java:100)
at com.caucho.server.http.request.parsequery(request.java:1352)
at com.caucho.server.http.request.getparametervalues(request.java:1449)
at com.caucho.server.http.request.getparameter(request.java:1459)
at com.chedong.weblucene.characterencodingtest.doget(unknown source)
at javax.servlet.http.httpservlet.service(httpservlet.java:126)
at javax.servlet.http.httpservlet.service(httpservlet.java:103)
at com.caucho.server.http.filterchainservlet.dofilter(filterchainservlet.java:96)
at com.caucho.server.http.invocation.service(invocation.java:315)
at com.caucho.server.http.cacheinvocation.service(cacheinvocation.java:135)
at com.caucho.server.http.httprequest.handlerequest(httprequest.java:246)
at com.caucho.server.http.httprequest.handleconnection(httprequest.java:163)
at com.caucho.server.tcpconnection.run(tcpconnection.java:139)
at java.lang.thread.run(thread.java:536)
附:
1.测试程序
1.测试程序
| package com.chedong.weblucene;
import org.apache.log4j.logger;
import java.io.ioexception;
import java.io.printwriter; import java.util.enumeration;
import javax.servlet.servletconfig;
import javax.servlet.servletexception; import javax.servlet.http.httpservlet; import javax.servlet.http.httpservletrequest; import javax.servlet.http.httpservletresponse; public class characterencodingtest extends httpservlet { //~ static fields/initializers --------------------------------------------- /** the global logger, it will be configured when the servlet loaded */
private static logger logger = logger.getlogger(characterencodingtest.class .getname() ); //~ methods ----------------------------------------------------------------
public void destroy() {
super.destroy(); } protected void doget(httpservletrequest request,
httpservletresponse response ) throws ioexception, servletexception { //request.setcharacterencoding("iso-8859-1");
response.setcontenttype("text/html;charset=utf-8");
printwriter out = response.getwriter(); out.println(""); out.println(""); out.println(""); out.println(""); out.println(""); system.out.println("i get a request, with encoding: " + request.getcharacterencoding() + "
"); string q = "empty";
if(request.getparameter("q") != null) { string ie = request.getparameter("ie"); if(ie == null) { ie = "gb2312"; } q = request.getparameter("q"); system.out.println("raw q:" + q + "."); q = new string(q.getbytes("iso-8859-1"), ie); system.out.println("parsed q with(" + ie + "):" + q); } else { q = "null"; enumeration params = request.getparameternames(); while(params.hasmoreelements()) { string param = (string) params.nextelement(); system.out.println("param: " + param + "-" + request.getparameter(param) + ". "); } }
本文关键:WebLucene 实现类似于Google 的多编码支持
相关方案
|