Showing posts with label Solr. Show all posts
Showing posts with label Solr. Show all posts

Salmon Run: Using Lucene Similarity in Item-Item Recommenders



By default, Lucene stores document vectors keyed by terms, but can be configured to store term vectors by setting the field attribute TermVector.YES. In case of text documents, words (or terms) are the features which are used to compute similarity between documents. I am using the same dataset as last week, where movies (items) correspond to documents and movie tags correspond to the words. So we build a movie "document" by preprocessing the tags to form individual tokens and concatenating them into a tags field in the index.

Find Movies Similar to given Movie: This is just content based filtering and is implemented as a simple MLT query. Given the itemID, we lookup the docID of the source movie, then get the top N movies that are most like it. We then return a List of tuples of docIDs that are similar and their similarities (scores), except the original docID.

Predict a User's Rating for a Movie: This is the prediction functionality of an item-item CF recommender. The prediction is based on how the user has rated other movies similar to this one. Otherwise, we calculate the average weighted sum of the ratings of already rated items, where the weights are the similarities between the target item and this item. If the movie is already rated, we just return the rating. Similarity between two items are calculated using the MLT query using a simplifying assumption - a target item outside the item neighborhood has 0 similarity with the source item. If we did not use this assumption, we would have to use approaches such as the TermFreqVector API for Lucene 3.x or the Fields API for Lucene 4.x to compute individual doc-doc similarities.

Recommend Movies to a User: This is topN recommender functionality of an item-item CF. We recommend movies that are similar to ones the user has already rated, weighted by the similarity between this item and the rated item. We use the algorithm outlined in Mahout in Action, § 4.4.1, detailed below. The 3rd and 4th lines in the algorithm is essentially the rating prediction task we described above. Essentially, we calculate the prediction for all items not rated so far by the user, and return them sorted by descending order of predicted rating.

Read full article from Salmon Run: Using Lucene Similarity in Item-Item Recommenders

DocValues - Solr Wiki



DocValues - Solr Wiki
With a search engine you typically build an inverted index (indexed="true") for a field: where values point to documents. DocValues is a way to build a forward index (docValues="true") so that documents point to values.
  1. What docvalues are:
    • NRT-compatible: These are per-segment datastructures built at index-time and designed to be efficient for the use case where data is changing rapidly.
    • Basic query/filter support: You can do basic term, range, etc queries on docvalues fields without also indexing them, but these are constant-score only and typically slower. If you care about performance and scoring, index the field too.
    • Better compression than fieldcache: Docvalues fields compress better than fieldcache, and "insanity" is impossible.
    • Able to store data outside of heap memory: You can specify a different docValuesFormat on the fieldType (docValuesFormat="Disk") to only load minimal data on the heap, keeping other data structures on disk.
  2. What docvalues are not:
    • Not a replacement for stored fields: These are unrelated to stored fields in every way and instead datastructures for search (sort/facet/group/join/scoring).
    • Not a huge improvement for a static index: If you have a completely static index, docvalues won't seem very interesting to you. On the other hand if you are fighting the fieldcache, read on.
    • Not for the risk-averse: The integration with Solr is very new and probably still has some exciting bugs!
SORTED: a single-valued per-document string type. This is like having a large String[] array for the whole index, but with an additional level of indirection. Each unique value is assigned a term number that represents its ordinal value. So each document really stores a compressed integer, and separately there is a "dictionary" mapping these term numbers back to term values.
SORTED_SET
    • alue "aardvark" will be assigned ordinal 0, "beaver" 1, and "cat" 2, creating these two data structures:
             doc[0] = [0, 1, 2]
             doc[1] = []
             doc[2] = [2]
      
             term[0] = "aardvark"
             term[1] = "beaver"
             term[2] = "cat"
  1. BINARY: a single-valued per-document byte[] array. This can be used for encoding custom per-document datastructures.
Read full article from DocValues - Solr Wiki

DocValues - Apache Solr Reference Guide - Apache Software Foundation



DocValues - Apache Solr Reference Guide - Apache Software Foundation
The standard way that Solr builds the index is with an inverted index. This style builds a list of terms found in all the documents in the index and next to each term is a list of documents that the term appears in (as well as how many times the term appears in that document). This makes search very fast - since users search by terms, having a ready list of term-to-document values makes the query process faster.
For other features that we now commonly associate with search, such as sorting, faceting, and highlighting, this approach is not very efficient. The faceting engine, for example, must look up each term that appears in each document that will make up the result set and pull the document IDs in order to build the facet list. In Solr, this is maintained in memory, and can be slow to load (depending on the number of documents, terms, etc.).
In Lucene 4.0, a new approach was introduced. DocValue fields are now column-oriented fields with a document-to-value mapping built at index time. This approach promises to relieve some of the memory requirements of the fieldCache and make lookups for faceting, sorting, and grouping much faster.

<field name="manu_exact" type="string" indexed="false" stored="false" docValues="true" />

DocValues are only available for specific field types. The types chosen determine the underlying Lucene docValue type that will be used. The available Solr field types are:
  • String fields of type StrField.
    • If the field is single-valued (i.e., multi-valued is false), Lucene will use the SORTED type.
    • If the field is multi-valued, Lucene will use the SORTED_SET type.
  • Any Trie* fields.
    • If the field is single-valued (i.e., multi-valued is false), Lucene will use the NUMERIC type.
    • If the field is multi-valued, Lucene will use the SORTED_SET type.
  • UUID fields
The default implementation employs a mixture of loading some things into memory and keeping some on disk.
<fieldType name="string_in_mem_dv" class="solr.StrField" docValues="true" docValuesFormat="Memory" />

Lucene index back-compatibility is only supported for the default codec. If you choose to customize the docValuesFormat in your schema.xml, upgrading to a future version of Solr may require you to either switch back to the default codec and optimize your index to rewrite it into the default codec before upgrading, or re-build your entire index from scratch after upgrading.

Read full article from DocValues - Apache Solr Reference Guide - Apache Software Foundation

Fun with DocValues in Solr 4.2 - Lucidworks



Fun with DocValues in Solr 4.2 - Lucidworks
The “dvd” and “dvm” files are the DocValues, “tim” and “tip” are the terms index and dictionary, and “fdx” and “fdt” are the stored fields. You can look up what the rest of those files are in the Lucene documentation. 

Without going any further, we can see that the DocValues are much more compact than the stored fields and the term index just by looking at the file sizes (recall that we are storing each of our fields as stored, indexed, and docValues separately). Since values for a single field are stored contiguously, very efficient packing algorithms can be used.
https://gist.github.com/mumrah/5265594
The performance differences here depend largely on the number of unique values for the field and the field type. The biggest difference in loading time is the field “word_idx”, which take twice as long to load from the inverted index and uses three times as much memory. For repeated access to the same field, the inverted index performs better due to internal Lucene caching (also the reason for higher memory consumption). In all cases, DocValues consume less memory during loading and after garbage collection.

DocValues have many potential uses. As we have seen from our little experiment, they are less memory hungry than indexed field and typically faster to load. If you are in a low-memory environment, or you don’t need to index a field, DocValues are perfect for faceting/grouping/filtering/sorting. They also have the potential for increasing the number of fields you can facet/group/filter/sort on without increasing your memory requirements.

Read full article from Fun with DocValues in Solr 4.2 - Lucidworks

Comparing Document Classification Functions of Lucene and Mahout | soleami | Visualize the needs of your visitors.



Comparing Document Classification Functions of Lucene and Mahout | soleami | Visualize the needs of your visitors.
Lucene implements Naive Bayes and k-NN rule classifiers. The trunk equivalent to Lucene 5, the next major releases, implements boolean (2-class) classification perceptron in addition to these two. We use Lucene 4.6.1, the most recent version at the time of writing, to perform document classification with Naive Bayes and k-NN rule.
You need to have IndexReader with prepared index open and specify it as the first argument of the train() method because Classifier uses index as learning data. Also, set the Lucene field name that has text, which is tokenized and indexed, as the second argument of train() method. In addition, set the Lucene field that has document category as the third argument of train() method. In the same manner, set a Lucene Analyzer to the fourth argument and Query to the fifth argument. Analyzer then specifies Analyzer that is used to classify unknown document (In my personal opinion, this is a bit complicated and should use them as arguments for after-mentioned assignClass() method instead) . While Query is used to narrow down documents that are used for learning, null is used if there’s no need to do so. The train() method has 2 more varieties that have different arguments but I will skip the explanation for now.
Use unknown document in the String type as an argument to call the assignClass() method after you call train() of Classifier interface to obtain the result of classification. Classifier is an interface that uses Java Generics, and the ClassificationResult class that uses type variable T is the returned value of assignClass().
Calling the getAssignedClass() method of ClassificationResult gives you a classification result of the type T.
Note that Lucene’s classifier is unique in that the train() method does little work while the assignClass() does most of the work. This is where it is very different from the other commonly used machine learning software. In the learning phase of commonly used machine learning software, a model file is created by learning corpus according to a selected machine learning algorithm (This is where the most time/effort is put into. As Mahout is based on Hadoop, it uses MapReduce to try to reduce the time required here). And in the classification phase, an unknown document is classified by referring to a previously created model file. This phase usually requires little resource.
As Lucene uses an index as a model file, train() method, which is a learning phase, does almost nothing here (Its learning completes as soon as index is created). Lucene’s index, however, is optimized to perform high-speed keyword search and is not in an appropriate format for document classification model file. Therefore, here we do document classification by searching index with the assignClass() method that is a classification phase. Contrary to commonly used machine learning software, Lucene’s classifier requires very high computing power in the classification phase. For sites mainly focused on searching, this function that enables document classification should be appealing as they can create indexes without additional cost.

SimpleNaiveBayesClassifier is the first implement class of Classifier interface. As you can see from the name, it’s a Naive Bayes classifier. Naive Bayes classification finds c where conditional probability P(c|d), the probability of class being c in document d, becomes the highest. Here you use Bayes’ theorem to do deformation of P(c|d) but you need to find P(c)P(d|c) to calculate class c with the highest probability. While you usually calculate logarithm to avoid underflow, the assignClass() method of SimpleNaiveBayesClassifier repeats this calculation as many times as the number of classes to perform MLE (maximum likelihood estimation).
Using Lucene KNearestNeighborClassifier
Another implement class for Classifier is KNearestNeighborClassifier. KNearestNeighborClassifier specifies k, which is no less than 1, in an argument for constructor to create an instance. You can use the program exactly the same as one for SimpleNaiveBayesClassifier. Only you need to do is to replace the portion that is creating an instance for SimpleNaiveBayesClassifier with KNearestNeighborClassifier.
The assignClass() method does all the work for KNearestNeighborClassifier as well in the same manner described before but one interesting point is that it is using Lucene MoreLikeThis. MoreLikeThis is a tool that sees document to become criteria as a query and performs search. With this, you can find documents that are similar to the ones to be criteria. KNearestNeighborClassifier uses MoreLikeThis to “k” number of documents that are most similar to the unknown document passed to the assignClass() method. Then, the majority rule is applied to that k number of documents to determine the document category of unknown document.
Executing the same program as KNearestNeighborClassifier will display the following when k=1.

In this article, we used the same corpus to do document classification of the both Lucene and Mahout to compare their results. The accuracy rate seems to be higher for Mahout but, as already stated, its learning data classification use not all word but only top 2,000 important words in the body field. On the other hand, Lucene’s classifier, which accuracy rate was only 70%, uses the all words in body field. Lucene will be able to pass the 90% accuracy rate if you have a field to hold only the words reviewed specially for document classification. It may also be a good idea to create another Classifier implement class for train() method that has such function.
I should add that the accuracy rate goes down to around 80% when you do not use test data for learning but test it as real unknown data.
Read full article from Comparing Document Classification Functions of Lucene and Mahout | soleami | Visualize the needs of your visitors.

Lucene 4 is Super Convenient for Developing NLP Tools



Lucene 4 is Super Convenient for Developing NLP Tools
Lucene 4.0 classes that I used for developing this system are as follows:
  • IndexSearcher, TermQuery, TopDocs
    This system calculates similarities of synonym candidates that consist of nouns extracted from keywords and their descriptions. The system determines that the candidate is a synonym of keyword if similarity is bigger than a threshold value and output it to a CSV file.
    But how I calculate the similarity of a keyword and its synonym candidate. This system determines the similarity by calculating the similarity of keyword description Aa and dictionary entry description set {Ab} that are written using synonym candidates.
    Thus, I have to find {Ab} where I used classes such as IndexSearcher, TermQuery, and TopDocsto to search description field using synonym candidate.
  • PriorityQueue
    Next, I have to pick out “feature word” from Aa and {Ab} to calculate similarity of the two. In order to do so, I select N most important words to structure feature vector. Here, I use TF*IDF of the target word as their degree of importance. See the above SlideShare for the detail. Here, I use PriorityQueue to select “N most important words”
  • DocsEnum, TotalHitCountCollector
    I used TF*IDF to calculate weight to extract the above feature word and used DocsEnum.freq() to obtain TF. docFreq (number of articles including synonym candidate), which is a required parameter to obtain IDF, has been calculated by passing TotalHitCountCollector to the search() method of IndexSearcher.
  • Terms, TermsEnum
    I use these classes to search “description” field for synonym candidates.
These are usage examples for Lucene 4.0 on this system. I also believe Lucene will be a great help for NLP tool developers as well. For lexical knowledge obtention task using Bootstrap, for example, I can use a cycle (1: pattern extraction, 2: pattern selection, 3: instance extraction, 4: instance selection) to obtain knowledge from a small number of seed instances. I believe that you can replace pattern extraction and instance extraction with a simple search task if you use Lucene for these tasks.
Please read full article from Lucene 4 is Super Convenient for Developing NLP Tools

Text categorization with Lucene and Solr



Text categorization with Lucene and Solr
Let the algorithm assign one or more labels (classes) to some item given some previous knowledge
l Spam filter
l Tagging system
l Digit recognition system
l Text categorization 

l Lucene already has a lot of features for common information retrieval needs
l Postings
l Term vectors
l Statistics
l Positions
l TF / IDF
l maybe Payloads
l etc.
l We may avoid bringing in new components
to do classification just leveraging what we
get for free from Lucene

l Lucene has so many features stored you can take advantage of for free
l Therefore writing the classification algorithm is relatively simple
l In many cases you’re just not adding anything to the architecture
l Your Lucene index was already there for searching l Lucene index is, to some extent, already a model which we just need to “query” with the proper algorithm
l And it is fast enough 

Classifier API
l Training
l void train(atomicReader, contentField, classField, analyzer) throws IOException


K Nearest neighbor classifier
l Fairly simple classification algorithm
l Given some new unseen item
l I search in my knowledge base the k items which are nearer to the new one
l I get the k classes assigned to the k nearest items
l I assign to the new item the class that is most frequent in the k returned items 

K Nearest neighbor classifier
l How can we do this in Lucene?
l We have VSM for representing documents as
vectors and eventually find distances
l Lucene MoreLikeThis module can do a lot for it
l Given a new document
l It’s represented as a MoreLikeThisQuery which filters
out too frequent words and helps on keeping only the
relevant tokens for finding the neighbors
l The query is executed returning only the first k results
l The result is then browsed in order to find the most
frequent class and that is then assigned with a score
of classFreq / k 

Naïve Bayes classifier
l Slightly more complicated
l Based on probabilities
l C = argmax( P(d|c) * P(c) )
l P(d|c) : likelihood
l P(c) : prior
l With some assumptions:
l bag of words assumption: positions don't matter
l conditional independence: the feature probabilities
are independent given a class


Things to consider - bootstrapping
l How are your first documents classified?
l Manually
l Categories are already there in the documents
l Someone is explicitly charged to do that (e.g. article
authors) at some point in time
l (semi) automatically
l Using some existing service / library
l With or without human supervision
l In either case the classifier needs something to
be fed with to be effective 

As specific search services
l A classification based more like this
l While indexing
l For automatic text categorization

Automatic text categorization
l Once a doc reaches Solr
l We can use the Lucene classifiers to automate assigning document’s category
l We can leverage existing Solr facilites for enhancing the indexing pipeline
l An UpdateChain can be decorated with one or more UpdateRequestProcessors

CategorizationUpdateRequestProcessorFactory
CategorizationUpdateRequestProcessor
l void processAdd(AddUpdateCommand
cmd) throws IOException
l String text = solrInputDocument.getFieldValue(“text”);
l String class = classifier.assignClass(text);
l solrInputDocument.addField(“cat”, class);
l Every now and then need to retrain to get latest stuff in the current index, but that can be done in the background without affecting performances 

CategorizationUpdateRequestProcessor
l Finer grained control
l Use automatic text categorization only if a value
does not exist for the “cat” field
l Add the classifier output class to the “cat” field only if it’s above a certain score 

Implement a MaxEnt Lucene based classifier
l which takes into account words correlation 

Please read full article from Text categorization with Lucene and Solr

Too Many Words Again! | HathiTrust Digital Library



Too Many Words Again! | HathiTrust Digital Library
Every indexed term that's loaded into RAM creates 4 objects (TermInfo,
Term, String, char[]), as you see in your profiler output.  And each
object has a number of fields, the header required by the JRE, GC
cost, etc. [2]

Even though the tii files took only  about  2.2 GB on disk  (about 750MB per index), once they are read into memory they occupy about 18 GB.
In Solr 1.4  and above there is a feature that lets you configure an “index divisor” for Solr.   If you set this to 2, then Solr will only load every other entry from the tii file into memory; thus halving memory use for the tii file representation in memory.  The downside is that once you have a file pointer into the tis file and seek to it, in the worst case you have to scan twice as many entries.[3]
Here is how Solr is configured to set it to 2:
<!-- To set   the termInfosIndexDivisor, do this: -->

<indexReaderFactory   class="org.apache.solr.core.StandardIndexReaderFactory">
  <int name="termInfosIndexDivisor">2</int>
</indexReaderFactory >
We upgraded the Solr on our test server to Solr 1.4.1 (in production we are currently using a pre-1.4 development release) and ran some tests with different settings.
We  set the termInfosIndexDivisor to 2 and then to 4 and ran a query against all 3 shards to cause the tii files to get loaded into memory.  We then ran jmap to get a histogram dump of the heap.  The table below shows the total memory use for the top 20 classes for each configuration including a base configuration where we don’t set the index divisor.

Base (current production config)
Index divisor =2
Index divisor =4
Total mem use for top 20 classes (GB)
17.9
9.6
6.1
We ran some preliminary tests and thus far have seen no significant impact in terms of response time for index divisors of 2, 4,8, and 16 with base memory use dropping as low as  a little over 1 GB [4].   We plan to do a few more tests to decide on which divisor to use and then to work on JVM tuning (We should be able to eliminate long stop-the-world collections with the proper settings.)  Once we get that done, we plan to upgrade our production Solrs to 1.4.1 and reduce the memory allocated to the JVM from 32 GB to some level possibly as low as 8 GB.  That will leave even more memory for the OS disk cache.  When we finish the tests and come up with a configuration and JVM settings we will report it in this blog.
Next up "Adventures in Garbage Collection" and "Why are there So Many Words?"
Read full article from Too Many Words Again! | HathiTrust Digital Library

Boosting Documents in Solr by Recency, Popularity, and User Preferences



Boosting Documents in Solr by Recency, Popularity, and User Preferences

Date published = DateUtils.round(item.getPublishedOnDate(),Calendar.HOUR);


FunctionQuery: Computes a value for each document
Ranking
Sorting

Use the recip function with the ms function:
q={!boost b=$recency v=$qq}&
 recency=recip(ms(NOW/HOUR,pubdate),3.16e-11,0.08,0.05)&
 qq=wine

Use edismax vs. dismax if possible:
 q=wine&
 boost=recip(ms(NOW/HOUR,pubdate),3.16e-11,0.08,0.05)

Recip is a highly tunable function
recip(x,m,a,b) implementing a / (m*x + b)
m = 3.16E-11 a= 0.08 b=0.05 x = Document Age

Boost should be a multiplier on the relevancy score 

{!boost b=} syntax confuses the spell checker so you need to use spellcheck.q to be explicit
q={!boost b=$recency v=$qq}&spellcheck.q=wine 

Bottom out the old age penalty using min:
min(recip(…), 0.20)

Not a one-size fits all solution – academic research focused on when to apply it 
Score based on number of unique views
Not known at indexing time
View count should be broken into time slots

fieldType name="externalPopularityScore"  
           keyField="id" 
           defVal="1" 
           stored="false" indexed="false" 
           class=”solr.ExternalFileField" 
           valType="pfloat"/>

<field name="popularity" 
       type="externalPopularityScore" />

For big, high traffic sites, use log analysis
Perfect problem for MapReduce
Take a look at Hive for analyzing large volumes of log data

Minimum popularity score is 1 (not zero) … up to 2 or more
1 + (0.4*recent + 0.3*lastWeek + 0.2*lastMonth …)

Watch out for spell checker “buildOnCommit”

Filtering By User Preferences
Easy approach is to build basic preference fields in to the index:
Content types of interest – content_type
High-level categories of interest - category
Source of interest – source

We had too many categories and sources that a user could enable / disable to use basic filtering
Custom SearchComponent with a connection to a JDBC DataSource

Connects to a database
Caches DocIdSet in a Solr FastLRUCache
Cached values marked as dirty using a simple timestamp passed in the request

Declared in solrconfig.xml:
  <searchComponent   
      class=“demo.solr.PreferencesComponent" 
      name=”pref">
    <str name="jdbcJndi">jdbc/solr</str>  
  </searchComponent>

Parameters passed in the query string:
pref.id = primary key in db
pref.mod = preferences modified on timestamp
So the Solr side knows the database has been updated
Use simple SQL queries to compute a list of disabled categories, feeds, and types
Lucene FieldCaches for category, source, type
Custom SearchComponent included in the list of components for edismax search handler
 <arr name="last-components">
      <str>pref</str>   
    </arr>
Use recip & ms functions to boost recent documents

Use ExternalFileField to load popularity scores calculated outside the index


Use a custom SearchComponent with a Solr FastLRUCache to filter documents using complex user preferences
Please read full article from Boosting Documents in Solr by Recency, Popularity, and User Preferences

Black Boxes: Monitoring Solr (JMX Edition) | AppNeta



Black Boxes: Monitoring Solr (JMX Edition) | AppNeta
Solr exposes hundreds of JMX metrics across dozens of categories, and efficient use of them can help you delve into Solr performance in a variety of ways. Some metrics are better for providing a high-level view of Solr’s overall workflow. The queryResultCache category, pictured above, provides a snapshot of how often your data was successfully cached, as well as how often cache entries had to be evicted due to insufficient space. Other metric categories are more granular and provide detail at the level of classes, or even objects. An update request will be routed to a different handler depending on whether the data was provided in XML, CSV, or JSON; each of these update handlers exposes metrics independently, like how long it has been running and the number of errors.
JMX metrics can even provide insight into advanced Solr use cases, like modifying result scoring to permit n-dimensional spatial searches or customizing results based on user data stored in Redis. Even without addingcustom JMX metrics, Solr will report enough data to allow you to separately track the effectiveness of these custom searches relative to more traditional queries.
After checking the metrics for that node’s active Searcher instance, you realize you didn’t set up Solr to warm the cache – it was starting off empty! Now you know to make a quick configuration change next time you spin up an instance so that the first users routed to it will have acceptable performance.

Purpose-built JMX monitoring tools like jconsole are great for browsing the available metrics to see what’s available, but they’re horrible for pulling out the ones you want in a hurry. They also allow ‘write’ operations like initiating garbage collection or clearing caches – definitely not something you want to give out to every developer!

On a day to day basis, it’s more common to read JMX metrics via automated, ‘read-only’ monitoring tools likeNagios, Ganglia, or AppNeta TraceView. These tools not only present a number of metrics at once, but they also generally let you filter down to a meaningful subset of the hundreds of lines exposed by Solr. On the other hand, “health check”-style metrics aren’t necessarily the only way to look the problem. Each request has a number of metrics it can generate, and bringing together these data sources in one application has some real advantages. Looking at an individual request can tell you exactly what went wrong, it’s often the context of JMX data that says why. Examining the concurrent host activity can disambiguate between whether a pause was due to a garbage collection event in the JVM or an overloaded document cache in Solr forcing additional disk access.

Read full article from Black Boxes: Monitoring Solr (JMX Edition) | AppNeta

New in Solr 4.8: Document Expiration - Lucidworks



New in Solr 4.8: Document Expiration – Lucidworks
The DocExpirationUpdateProcessorFactory provides two features related to the “expiration” of documents which can be used individually, or in combination:
  • Periodically delete documents from the index based on an expiration field
  • Computing expiration field values for documents from a “time to live” (TTL)

Auto-Delete Expired Documents

The biggest aspect of this Update Processor is it’s ability to automatically delete documents based on the values found in an “expiration date” field that you configure. This automatic deletion isn’t part of the normal Update Processor life cycle — it’s executed via a background timer process thread created by the Factory.
To use this automatic deletion feature, you must configure two options on the Factory:
  • expirationFieldName – The name of the expiration field to use
  • autoDeletePeriodSeconds – How often the factory’s timer should trigger a delete to remove the documents
For example, with the configuration below the DocExpirationUpdateProcessorFactory will create a timer thread that wakes up every 30 seconds. When the timer triggers, it will execute a deleteByQuerycommand to remove any documents with a value in the press_release_expiration_date field value that is in the past:
 <processor class="solr.processor.DocExpirationUpdateProcessorFactory">
   <int name="autoDeletePeriodSeconds">30</int>
   <str name="expirationFieldName">press_release_expiration_date</str>
 </processor>
After the deleteByQuery has been executed, a soft commit is also executed usingopenSearcher=true so that search results will no longer see the expired documents.
While the basic logic of “timer goes off, delete docs with expiration prior to NOW” was fairly simple and straight forward to add, a key aspect of making this work well was in a related issue (SOLR-5783) to ensure that the openSearcher=true doesn’t do anything unless there really are changes in the index. This means that you can configure autoDeletePeriodSeconds to very small values, and still rest easy that your search caches won’t get blown away every few seconds for no reason. The openSearcher=truesoft commits will only affect things if there really are changes in the index.

Compute Expiration Date from TTL

The second feature implemented by this Factory (and the key reason it’s implemented as anUpdateProcessorFactory) is to use “TTL” (Time To Live) values associated with documents to automatically generate an expiration date value to put in the expirationFieldName when documents are indexed.
By default, the DocExpirationUpdateProcessorFactory will look for a _ttl_ request parameter on update requests, as well as a _ttl_ field in each doc that is indexed in that request. If either exist, they will be parsed as Date Math Expressions relative to NOW and used to populate the expirationFieldName. The per-document _ttl_ field based values override the per-request _ttl_ parameter.
Both the request parameter and field names use for specifying TTL values can be overridden by configuringttlParamName & ttlFieldName on the DocExpirationUpdateProcessorFactory. They can also be completely disabled by configuring them as null. It’s also possible to use the TTL computation feature to generate expiration dates on documents, with out using the auto-deletion feature simply by not configuring the autoDeletePeriodSeconds option (so that the timer will never run).
For example, in the configuration below, the Factory will look for a time_to_live field in each document, and use that to compute an expiration value for the press_release_expiration_date field. No request parameters will be checked for a TTL override, and no automatic deletion will occur:
 <processor class="solr.processor.DocExpirationUpdateProcessorFactory">
   <str name="expirationFieldName">press_release_expiration_date</str>
   <null name="ttlParamName"/> <!-- ignore _ttl_ request param -->
   <str name="ttlFieldName">time_to_live</str>
   <!-- NOTE: autoDeletePeriodSeconds not specified, no automatic deletes -->
 </processor>
This sort of configuration may be handy if you only want to logically hide documents for search clients based on a per-document TTL using something like: fq=-press_release_expiration_date:[* TO NOW/DAY], but still retain the documents in the index for other search clients.

An In Depth Example

Let’s walk through a full example of both features by modifying the Solr 4.8 example solrconfig.xml to add the following update processor chain:
  <updateRequestProcessorChain default="true">
    <processor class="solr.TimestampUpdateProcessorFactory">
      <str name="fieldName">timestamp_dt</str>
    </processor>
    <processor class="solr.processor.DocExpirationUpdateProcessorFactory">
      <int name="autoDeletePeriodSeconds">30</int>
      <str name="ttlFieldName">time_to_live_s</str>
      <str name="expirationFieldName">expire_at_dt</str>
    </processor>
    <processor class="solr.FirstFieldValueUpdateProcessorFactory">
      <str name="fieldName">expire_at_dt</str>
    </processor>
    <processor class="solr.LogUpdateProcessorFactory" />
    <processor class="solr.RunUpdateProcessorFactory" />
  </updateRequestProcessorChain>
A few things to note about this chain:
  • It contains a simple TimestampUpdateProcessorFactory so that it will be easy to see when these documents were indexed in the query results I show below — but this is not needed forDocExpirationUpdateProcessorFactory to function
  • The DocExpirationUpdateProcessorFactory instance uses a autoDeletePeriodSeconds of 30 seconds and overrides the ttlFieldName – but the _ttl_ request param is still enabled
  • A FirstFieldValueUpdateProcessorFactory is configured on the expire_at_dt — this means that if a document is added with an explicit value in the expire_at_dt field, it will be used instead any value that might be added by the DocExpirationUpdateProcessorFactory using the _ttl_request param
Read full article from New in Solr 4.8: Document Expiration – Lucidworks

Solr 4.8 Features Overview



Solr 4.8 Features Overview

Complex Phrase Queries

The complexphrase query parser can produce phrase queries with embedded wildcards and boolean queries.
It works via multiple passes, parsing a query and then re-parsing any phrase queries for additional markup. At query execution time, span queries are generated to implement the complex phrase logic.
The simplest example is a phrase query containing a prefix query:
q={!complexphrase}"apple ip*"
This will match text with both “apple ipod” and “apple ipad”.
One can specify inOrder=false as a localParam to also match “ipod apple” and “ipad apple”.
q={!complexphrase inOrder=false}"apple ip*"
One can also specify a different default field to search with the df localParam:
q={!complexphrase df=name}"john* smith"
This will match both “john smith” and “johnathan smith” in the name field. Of course one could always specify the field directly in the query as well:
q={!complexphrase}name:"john* smith"
Phrase slop works to specify the proximity of the clauses. For example, the following would also match a name of “johnathan q smith”:
q={!complexphrase}name:"john* smith"~1
And of course we can throw in parens, OR clauses, and other complex logic as well:
q={!complexphrase}name:"(aaa OR (bbb* OR ccc)) ddd -eee (fff~1 OR ggg)" AND text:"nnn? (ooo OR ppp) -qqq www"~3
Indexing Child Documents in JSON
Named Config Sets

This is more in the “configuration” category of features. SolrCloud has always allowed multiple collections to share configuration, and now that capability has been brought to Solr’s non-cloud mode.
Since collections can be created or destroyed, we obviously don’t want shared configuration for these collections to be under the collection itself. The default location for config sets is in the “configsets” directory under the solr home (the example solr server currently doesn’t have this directory by default).

Stopwords and Synonyms REST API

Stopwords and Synonyms may now be managed via a REST API!
The new analysis filter types are ManagedStopFilterFactory and ManagedSynonymFilterFactory.
The example schema.xml now contains a field type that uses these new analysis filters:
<!-- A text type for English text where stopwords and synonyms are managed using the REST API -->
<fieldType name="managed_en" class="solr.TextField" positionIncrementGap="100">
  <analyzer>
    <tokenizer class="solr.StandardTokenizerFactory"/>
    <filter class="solr.ManagedStopFilterFactory" managed="english" />
    <filter class="solr.ManagedSynonymFilterFactory" managed="english" />
  </analyzer>
</fieldType>
To test this out, let’s also change the dynamic field *_en to use managed_en:
<dynamicField name="*_en"  type="managed_en"    indexed="true"  stored="true" multiValued="true"/>

Synonyms

After starting the example server, we can retrieve the current english synonyms:
[...]
    "managedMap":{
      "gb":["gib",
        "gigabyte"],
      "happy":["glad",
        "joyful"],
      "tv":["television"]}}}

Lets add a new synonym:
curl -XPUT "http://localhost:8983/solr/collection1/schema/analysis/synonyms/english" -H 'Content-type:application/json' --data-binary '{"mb":["MiB","megabyte"]}'

Before these changes are visible to the actual search or indexing code in Solr, we need to reload the Solr core:

And now we can do a query on a field that matches the dynamicField we set up and can see the results of the new synonym:
[...]
  "debug":{
    "rawquerystring":"foo_en:mb",
    "querystring":"foo_en:mb",
    "parsedquery":"(foo_en:megabyte foo_en:mib)/no_coord",
    "parsedquery_toString":"foo_en:megabyte foo_en:mib",

To delete the stopword we just added:

Stopwords

To retrieve the list of stopwords:
To add a new stopword:
curl -XPUT "http://localhost:8983/solr/collection1/schema/analysis/stopwords/english" -H 'Content-type:application/json' --data-binary '["foo"]'
To delete the stopword we just added:

Other changes

There have been numerous SolrCloud changes, including:
  • A new List collections and cluster status API which clients can use to read collection and shard information instead of reading data directly from ZooKeeper.
  • Some long running SolrCloud commands (like shard splitting) may now be run in “async” mode to avoid client timeouts
  • A new ADDREPLICA command in the Collections API
Other changes include:
  • Solr 4.8 now requires Java7!
  • RegexReplaceProcessorFactory now supports pattern capture group substitution in the replacement string.
  • A DocExpirationUpdateProcessorFactory that can mark documents based on a TTL (time-to-live) and periodically delete expired documents
Read full article from Solr 4.8 Features Overview

Labels

Algorithm (219) Lucene (130) LeetCode (97) Database (36) Data Structure (33) text mining (28) Solr (27) java (27) Mathematical Algorithm (26) Difficult Algorithm (25) Logic Thinking (23) Puzzles (23) Bit Algorithms (22) Math (21) List (20) Dynamic Programming (19) Linux (19) Tree (18) Machine Learning (15) EPI (11) Queue (11) Smart Algorithm (11) Operating System (9) Java Basic (8) Recursive Algorithm (8) Stack (8) Eclipse (7) Scala (7) Tika (7) J2EE (6) Monitoring (6) Trie (6) Concurrency (5) Geometry Algorithm (5) Greedy Algorithm (5) Mahout (5) MySQL (5) xpost (5) C (4) Interview (4) Vi (4) regular expression (4) to-do (4) C++ (3) Chrome (3) Divide and Conquer (3) Graph Algorithm (3) Permutation (3) Powershell (3) Random (3) Segment Tree (3) UIMA (3) Union-Find (3) Video (3) Virtualization (3) Windows (3) XML (3) Advanced Data Structure (2) Android (2) Bash (2) Classic Algorithm (2) Debugging (2) Design Pattern (2) Google (2) Hadoop (2) Java Collections (2) Markov Chains (2) Probabilities (2) Shell (2) Site (2) Web Development (2) Workplace (2) angularjs (2) .Net (1) Amazon Interview (1) Android Studio (1) Array (1) Boilerpipe (1) Book Notes (1) ChromeOS (1) Chromebook (1) Codility (1) Desgin (1) Design (1) Divide and Conqure (1) GAE (1) Google Interview (1) Great Stuff (1) Hash (1) High Tech Companies (1) Improving (1) LifeTips (1) Maven (1) Network (1) Performance (1) Programming (1) Resources (1) Sampling (1) Sed (1) Smart Thinking (1) Sort (1) Spark (1) Stanford NLP (1) System Design (1) Trove (1) VIP (1) tools (1)

Popular Posts