Class CandidateElementExtractor


  • public class CandidateElementExtractor
    extends Object
    This class extracts candidate elements from the DOM tree, based on the tags provided by the user. Elements can also be excluded.
    • Field Detail

      • LOG

        private static final org.slf4j.Logger LOG
      • crawlFrames

        private final boolean crawlFrames
      • excludeCrawlElements

        private final com.google.common.collect.ImmutableMultimap<String,​CrawlElement> excludeCrawlElements
      • includedCrawlElements

        private final com.google.common.collect.ImmutableList<CrawlElement> includedCrawlElements
      • clickOnce

        private final boolean clickOnce
      • randomizeElementsOrder

        private final boolean randomizeElementsOrder
      • ignoredFrameIdentifiers

        private final com.google.common.collect.ImmutableSortedSet<String> ignoredFrameIdentifiers
      • followExternalLinks

        private final boolean followExternalLinks
      • siteHostName

        private final String siteHostName
    • Constructor Detail

      • CandidateElementExtractor

        @Inject
        public CandidateElementExtractor​(ExtractorManager checker,
                                         EmbeddedBrowser browser,
                                         FormHandler formHandler,
                                         CrawljaxConfiguration config)
        Create a new CandidateElementExtractor.
        Parameters:
        checker - the ExtractorManager to use for marking handled elements and retrieve the EventableConditionChecker
        browser - the current browser instance used in the Crawler
        formHandler - the form handler.
        config - the checker used to determine if a certain frame must be ignored.
    • Method Detail

      • asMultiMap

        private com.google.common.collect.ImmutableMultimap<String,​CrawlElement> asMultiMap​(com.google.common.collect.ImmutableList<CrawlElement> elements)
      • extract

        public com.google.common.collect.ImmutableList<CandidateElement> extract​(StateVertex currentState)
                                                                          throws CrawljaxException
        This method extracts candidate elements from the current DOM tree in the browser, based on the crawl tags defined by the user.
        Parameters:
        currentState - the state in which this extract method is requested.
        Returns:
        a list of candidate elements that are not excluded.
        Throws:
        CrawljaxException - if the method fails.
      • isFrameIgnored

        private boolean isFrameIgnored​(String string)
      • getNodeListForTagElement

        private com.google.common.collect.ImmutableList<Element> getNodeListForTagElement​(Document dom,
                                                                                          CrawlElement crawlElement,
                                                                                          EventableConditionChecker eventableConditionChecker)
        Returns a list of Elements form the DOM tree, matching the tag element.
      • getFullXpathForGivenXpath

        private com.google.common.collect.ImmutableList<String> getFullXpathForGivenXpath​(Document dom,
                                                                                          EventableCondition eventableCondition)
      • addElement

        private void addElement​(Element element,
                                com.google.common.collect.ImmutableList.Builder<Element> builder,
                                CrawlElement crawlElement)
      • hrefShouldBeIgnored

        boolean hrefShouldBeIgnored​(Element element)
      • isExternal

        private boolean isExternal​(String href)
      • isFileForDownloading

        private boolean isFileForDownloading​(String href)
        Parameters:
        href - the string to check
        Returns:
        true if href has the pdf or ps pattern.
      • isExcluded

        private boolean isExcluded​(Document dom,
                                   Element element,
                                   EventableConditionChecker eventableConditionChecker)
        Returns:
        true if element should be excluded. Also when an ancestor of the given element is marked for exclusion, which allows for recursive exclusion of elements from candidates.
      • checkCrawlCondition

        public boolean checkCrawlCondition()