Using Lucene's new QueryParser framework in Solr

Sometime back, I described how I built (among other things) a custom Solr QParser plugin to handle Payload Term Queries. Looking back on this recently, I realized how lame it was - all it could handle were single Payload Term Queries, and a one level deep AND and OR combinations of these queries. More to the point, I discovered that I had to support queries of this form:

1	+concepts:123456 +(concepts:234567 concepts:345678 ...)

The original parsing code simply split up the query by whitespace, then by colon, and depending on whether the key was preceded by a "+" sign, either added it to the Boolean Query as an Occur.MUST or Occur.SHOULD. Obviously, this would not be able to parse the form of the query above.

Coencidentally, a few days ago, I was hunting around for something completely different on my laptop, and I came across the QueryParser Lucene contrib module that replaces the original Lucene JavaCC based QueryParser with a nice little framework that splits the query parsing into 3 phases - syntax parsing, query processing and query building. It has been available since Lucene 2.9.0, and on the version I am using (Lucene 2.9.3/Solr 1.4.1) both QueryParser implementations are supported.

In my case, my Payload Query syntax is identical to the Term Query syntax, so all I really needed to do was to return a PayloadTermQuery instead of a TermQuery in the query building phase. So all I needed to do to build a robust Payload QueryParser was to just implement a custom QueryBuilder and call it from within this framework.

There is not much documentation available on how to use the framework though, apart from the Javadocs, and the advice in there is to take a look at the StandardQueryParser and use that as a template to design your own. So thats what I did. I ended up building a few more classes in order to integrate it into my custom QParser plugin, but it was really quite simple.

Here is the updated code for my QParser plugin. Apart from this code change, all I had to do was add the lucene-queryparser-2.9.3.jar to the Solr classpath. There is no change in its configuration and the associated Solr request handler I used it from - both these are described in my previous post I referred to above.

// $Source: src/java/org/apache/solr/search/ext/PayloadQParserPlugin.java

package org.apache.solr.search.ext;

import org.apache.lucene.index.Term;

import org.apache.lucene.queryParser.ParseException;

import org.apache.lucene.queryParser.core.QueryNodeException;

import org.apache.lucene.queryParser.core.QueryParserHelper;

import org.apache.lucene.queryParser.core.nodes.FieldQueryNode;

import org.apache.lucene.queryParser.core.nodes.QueryNode;

import org.apache.lucene.queryParser.standard.builders.StandardQueryBuilder;

import org.apache.lucene.queryParser.standard.builders.StandardQueryTreeBuilder;

import org.apache.lucene.queryParser.standard.config.StandardQueryConfigHandler;

import org.apache.lucene.queryParser.standard.parser.StandardSyntaxParser;

import org.apache.lucene.queryParser.standard.processors.StandardQueryNodeProcessorPipeline;

import org.apache.lucene.search.Query;

import org.apache.lucene.search.payloads.AveragePayloadFunction;

import org.apache.lucene.search.payloads.PayloadTermQuery;

import org.apache.solr.common.params.SolrParams;

import org.apache.solr.common.util.NamedList;

import org.apache.solr.request.SolrQueryRequest;

import org.apache.solr.search.QParser;

import org.apache.solr.search.QParserPlugin;

/**

 * Parser plugin to parse payload queries.

 */

public class PayloadQParserPlugin extends QParserPlugin {

  @Override

  public QParser createParser(String qstr, SolrParams localParams,

      SolrParams params, SolrQueryRequest req) {

    return new PayloadQParser(qstr, localParams, params, req);

  }

  public void init(NamedList args) {

    // do nothing

  }

}

class PayloadQParser extends QParser {

  public PayloadQParser(String qstr, SolrParams localParams,

      SolrParams params, SolrQueryRequest req) {

    super(qstr, localParams, params, req);

  }

  @Override

  public Query parse() throws ParseException {

    PayloadQueryParser parser = new PayloadQueryParser();

    try {

      Query q = (Query) parser.parse(qstr, "concepts");

      return q;

    } catch (QueryNodeException e) {

      throw new ParseException(e.getMessage());

    }

  }

}

class PayloadQueryParser extends QueryParserHelper {

  public PayloadQueryParser() {

    super(new StandardQueryConfigHandler(), new StandardSyntaxParser(),

      new StandardQueryNodeProcessorPipeline(null),

      new PayloadQueryTreeBuilder());

  }

}

class PayloadQueryTreeBuilder extends StandardQueryTreeBuilder {

  public PayloadQueryTreeBuilder() {

    super();

    setBuilder(FieldQueryNode.class, new PayloadQueryNodeBuilder());

  }

}

class PayloadQueryNodeBuilder implements StandardQueryBuilder {

  @Override

  public PayloadTermQuery build(QueryNode queryNode) throws QueryNodeException {

    FieldQueryNode node = (FieldQueryNode) queryNode;

    return new PayloadTermQuery(

      new Term(node.getFieldAsString(), node.getTextAsString()),

      new AveragePayloadFunction(), false);

  }

}

As you can see, in my QParser.parse() method, I instantiate PayloadQueryParser, which is a subclass of QueryParserHelper. I reuse the same constructor code as StandardQueryParser (another subclass of QueryParserHelper and my template), except I pass in a custom QueryBuilder - the PayloadQueryTreeBuilder. The PayloadQueryTreeBuilder subclasses StandardQueryTreeBuilder, except it redefines what builder to use for FieldQueryNode types - the StandardQueryTreeBuilder is sort of a factory and delegates to the appropriate QueryBuilder depending on the type of the node. Finally, the PayloadQueryNodeBuilder implements the StandardQueryBuilder (similar to the FieldQueryNodeBuilder), and redefines the build() method to produce a PayloadTermQuery instead of a TermQuery as FieldQueryNodeBuilder does.

And thats pretty much it. I tested this by hitting the /concept-search URL and verified that the queries are correctly parsed and returned by printing the queries in the log.

Hopefully this post was useful, if for nothing else that people find out about the new QueryParser framework and begin to use it. The customization I did here is pretty trivial in terms of code, but it saved me a lot of work.

http://sujitpal.blogspot.com/2011/03/using-lucenes-new-queryparser-framework.html

Using Lucene's new QueryParser framework in Solr的更多相关文章

Lucene自定义扩展QueryParser
Lucene版本:4.10.2 在使用lucene的时候,不可避免的需要扩展lucene的相关功能来实现业务的需要,比如搜索时,需要在满足一个特定范围内的document进行搜索,如年龄在20和30岁 ...
solr调用lucene底层实现倒排索引源码解析
1.什么是Lucene? 作为一个开放源代码项目,Lucene从问世之后,引发了开放源代码社群的巨大反响,程序员们不仅使用它构建具体的全文检索应用,而且将之集成到各种系统软件中去,以及构建Web应用, ...
Lucene&Solr框架之第二篇
2.1.开发环境准备 2.1.1.数据库jar包我们这里可以尝试着从数据库中采集数据,因此需要连接数据库,我们一直用MySQL,所以这里需要MySQL的jar包 2.1.2.MyBatis的jar包 ...
apache lucene solr 官网历史版本下载地址
官网上一般只提供最新版本的下载,下面两个链接为所有历史版本的下载地址: lucene地址:archive.apache.org/dist/lucene/java/ solr地址:archive.apa ...
Lucene/Solr开发经验
1.开篇语2.概述3.渊源4.初识Solr5.Solr的安装6.Solr分词顺序7.Solr中文应用的一个实例8.Solr的检索运算符 [开篇语]按照惯例应该写一篇技术文章了,这次结合Lucene/S ...
在Lucene或Solr中实现高亮的策略
一:功能背景近期要做个高亮的搜索需求,曾经也搞过.所以没啥难度.仅仅只是原来用的是Lucene,如今要换成Solr而已,在Lucene4.x的时候,散仙在曾经的文章中也分析过怎样在搜索的时候实现高亮 ...
lucene和solr
我们为什么要用solr呢? 1.solr已经将整个索引操作功能封装好了的搜索引擎系统(企业级搜索引擎产品) 2.solr可以部署到单独的服务器上(WEB服务),它可以提供服务,我们的业务系统就只要发送 ...
Elasticsearch、Solr、Lucene、Hermes区别
Elasticsearch简介 Elasticsearch是一个实时分布式搜索和分析引擎.它让你以前所未有的速度处理大数据成为可能.它用于全文搜索.结构化搜索.分析以及将这三者混合使用:维基百科使用E ...
Importing/Indexing database (MySQL or SQL Server) in Solr using Data Import Handler--转载
原文地址:https://gist.github.com/maxivak/3e3ee1fca32f3949f052 Install Solr download and install Solr fro ...

随机推荐

VNC跨平台远程桌面的安装与使用
1.安装:yum install tigervnc-server -y 2.设置自启动: chkconfig vncserver on 3.配置文件:vim /etc/sysconfig/vncser ...
linux rar安装
1.wget http://www.rarsoft.com/rar_CN/rarlinux-3.9.3.tar.gz 2.tar 3.make && make install; 4.需 ...
关于java中word转html
http://www.dewen.org/q/5820/java%E4%B8%AD%E6%80%8E%E4%B9%88%E5%B0%86word%E6%96%87%E6%A1%A3%E8%BD%ACh ...
你应该使用 Django admin 的 9 个理由（转）
你应该使用 Django admin 的 9 个理由 “问题是,我问到的每个人都持反对意见,他们认为 admin 只限于超级用户,很不灵活并且是难以定制.”—来自 Reddit 的 andybak 我 ...
UNITY 模型与动画优化选项
1,RIG: Optimze Game Objects,[默认是没勾选的] 效果:将骨骼层级从模型中移除,放到动画控制器中,这样性能提高明显.实测中发现原来瞬间加载5个场景角色有点延迟,采用此选项后流 ...
HAVING COUNT(*) > 1的用法和理解
HAVING COUNT(*) > 1的用法和理解作用是保留包含多行的组. SELECT class.STUDENT_CODE FROM crm_class_schedule class GR ...
springboot 中使用thymeleaf
Spring Boot支持FreeMarker.Groovy.Thymeleaf和Mustache四种模板解析引擎,官方推荐使用Thymeleaf. spring-boot-starter-thyme ...
【 Makefile 编程基础之…
本站文章均为李华明Himi原创,转载务必在明显处注明: 转载自[黑米GameDev街区] 原文链接: http://www.himigame.com/gcc-makefile/766.html 概述: ...
LeetCode之位操作题java
191. Number of 1 Bits Total Accepted: 87985 Total Submissions: 234407 Difficulty: Easy Write a funct ...
MySQL内置功能之事务、函数和流程控制
主要内容: 一.事务二.函数三.流程控制 1️⃣ 事务一.何谓事务? 事务用于将某些操作的多个SQL作为原子性操作,一旦有某一个出现错误,即可回滚到原来的状态,从而保证数据库数据完整性. # ...

Using Lucene's new QueryParser framework in Solr

Using Lucene's new QueryParser framework in Solr的更多相关文章

随机推荐

热门专题