Class MetadataIndexer

java.lang.Object
org.apache.nutch.indexer.metadata.MetadataIndexer
All Implemented Interfaces:
Configurable, IndexingFilter, Pluggable

public class MetadataIndexer extends Object implements IndexingFilter
Indexer which can be configured to extract metadata from the crawldb, parse metadata or content metadata. You can specify the properties "index.db.md", "index.parse.md" or "index.content.md" who's values are comma-delimited key1,key2,key3.
  • Constructor Details

    • MetadataIndexer

      public MetadataIndexer()
  • Method Details

    • filter

      public NutchDocument filter(NutchDocument doc, Parse parse, Text url, CrawlDatum datum, Inlinks inlinks) throws IndexingException
      Description copied from interface: IndexingFilter
      Adds fields or otherwise modifies the document that will be indexed for a parse. Unwanted documents can be removed from indexing by returning a null value.
      Specified by:
      filter in interface IndexingFilter
      Parameters:
      doc - document instance for collecting fields
      parse - parse data instance
      url - page url
      datum - crawl datum for the page (fetch datum from segment containing fetch status and fetch time)
      inlinks - page inlinks
      Returns:
      modified (or a new) document instance, or null (meaning the document should be discarded)
      Throws:
      IndexingException - if an error occurs during during filtering
    • add

      protected void add(NutchDocument doc, String key, String value)
    • setConf

      public void setConf(Configuration conf)
      Specified by:
      setConf in interface Configurable
    • getConf

      public Configuration getConf()
      Specified by:
      getConf in interface Configurable