Package com.crawljax.util
Class DomUtils
- java.lang.Object
-
- com.crawljax.util.DomUtils
-
public final class DomUtils extends Object
Utility class that contains a number of helper functions used by Crawljax and some plugins.
-
-
Field Summary
Fields Modifier and Type Field Description (package private) static intBASE_LENGTHprivate static org.slf4j.LoggerLOGGERprivate static intTEXT_CUTOFF
-
Constructor Summary
Constructors Modifier Constructor Description privateDomUtils()
-
Method Summary
All Methods Static Methods Concrete Methods Modifier and Type Method Description private static voidaddAttributesToString(com.google.common.collect.ImmutableSet<String> exclude, StringBuilder buffer, NamedNodeMap attributes)static StringaddFolderSlashIfNeeded(String folderName)Adds a slash to a path if it doesn't end with a slash.static DocumentasDocument(String html)transforms a string into a Document object.static booleancontains(Node parent, Node child)returns true if parent and child are the same nodeprivate static StringfilterAttributes(String html)Filters attributes from the HTML string.static Map<String,Set<String>>getAllAttributes(Document dom, Map<String,Set<String>> map, Set<String> filterSet)static StringgetAllElementAttributes(Element element)static List<String>getAllImageURLs(Document document)static NodeListgetAllLeafNodes(Document document)static NodeListgetAllSubtreeNodes(Node node)static StringgetAttributeFromElement(Node element, String attribute)static List<org.custommonkey.xmlunit.Difference>getDifferences(String controlDom, String testDom)Get differences between DOMs.static List<org.custommonkey.xmlunit.Difference>getDifferences(String controlDom, String testDom, List<String> ignoreAttributes)Get differences between DOMs.static DocumentgetDocumentNoBalance(String html)static byte[]getDocumentToByteArray(Document dom)Serialize the Document object.static StringgetDocumentToString(Document dom)static StringgetDOMContent(String dom)static StringgetDOMContent(Document document)static StringgetDOMWithoutContent(String dom)static StringgetDOMWithoutContent(Document document)static StringgetDomWithoutHead(String html)static StringgetElementAttributes(Element element, com.google.common.collect.ImmutableSet<String> exclude)static ElementgetElementByXpath(Document dom, String xpath)static List<Node>getElementsByTagName(Document document, String tagName)static StringgetElementString(Element element)private static StringgetFileNameInPath(String path)Returns the filename in a path.static StringgetFrameIdentification(Element frame)static intgetNumLeafNodes(Node node)static StringgetStrippedDom(String fullDom)static StringgetTemplateAsString(String fileName)Retrieves the content of the filename.static StringgetTextContent(Document document, boolean individualTokens)To get all the textual content in the domstatic List<String>getTextTokens(Document document)static StringgetTextValue(Element element)Returns the text value of an element (title, alt or contents).static DocumentremoveComments(Document document)static DocumentremoveElementsUnderXpath(Document dom, String xpath)Removes all the given tags from the document.static DocumentremoveHead(Document document)static DocumentremoveHiddenInputs(Document dom)static StringremoveNewLines(String html)Removes newlines from a string.static DocumentremoveScriptTags(Document dom)Removes all the <SCRIPT/> tags from the document.static DocumentremoveTags(Document dom, String tagName)Removes all the given tags from the document.static StringreplaceString(String string, String regex, String replace)private static StringtoUniformDOM(String html)static voidwriteDocumentToFile(Document document, String filePathname, String method, int indent)Write the document object to a file.
-
-
-
Field Detail
-
BASE_LENGTH
static final int BASE_LENGTH
- See Also:
- Constant Field Values
-
LOGGER
private static final org.slf4j.Logger LOGGER
-
TEXT_CUTOFF
private static final int TEXT_CUTOFF
- See Also:
- Constant Field Values
-
-
Method Detail
-
asDocument
public static Document asDocument(String html) throws IOException
transforms a string into a Document object. TODO This needs more optimizations. As it seems the getDocument is called way too much times causing a lot of parsing which is slow and not necessary.- Parameters:
html- the HTML string.- Returns:
- The DOM Document version of the HTML string.
- Throws:
IOException- if an IO failure occurs.
-
getDocumentNoBalance
public static Document getDocumentNoBalance(String html) throws SAXException, IOException
- Parameters:
html- the HTML string.- Returns:
- a Document object made from the HTML string.
- Throws:
SAXException- if an exception occurs while parsing the HTML string.IOException- if an IO failure occurs.
-
getAllElementAttributes
public static String getAllElementAttributes(Element element)
- Parameters:
element- The DOM Element.- Returns:
- A string representation of all the element's attributes.
-
getElementAttributes
public static String getElementAttributes(Element element, com.google.common.collect.ImmutableSet<String> exclude)
- Parameters:
element- The DOM Element.exclude- the list of exclude strings.- Returns:
- A string representation of the element's attributes excluding exclude.
-
addAttributesToString
private static void addAttributesToString(com.google.common.collect.ImmutableSet<String> exclude, StringBuilder buffer, NamedNodeMap attributes)
-
getElementString
public static String getElementString(Element element)
- Parameters:
element- the element.- Returns:
- a string representation of the element including its attributes.
-
getElementByXpath
public static Element getElementByXpath(Document dom, String xpath) throws XPathExpressionException
- Parameters:
dom- the DOM document.xpath- the xpath.- Returns:
- The element found on DOM having the xpath position.
- Throws:
XPathExpressionException- if the xpath fails.
-
removeScriptTags
public static Document removeScriptTags(Document dom)
Removes all the <SCRIPT/> tags from the document.- Parameters:
dom- the document object.- Returns:
- the changed dom.
-
removeElementsUnderXpath
public static Document removeElementsUnderXpath(Document dom, String xpath)
Removes all the given tags from the document.- Parameters:
dom- the document object.xpath- the tag name, examples: script, style, meta- Returns:
- the changed dom.
-
removeTags
public static Document removeTags(Document dom, String tagName)
Removes all the given tags from the document.- Parameters:
dom- the document object.tagName- the tag name, examples: script, style, meta- Returns:
- the changed dom.
-
getDocumentToString
public static String getDocumentToString(Document dom)
- Parameters:
dom- the DOM document.- Returns:
- a string representation of the DOM.
-
getDocumentToByteArray
public static byte[] getDocumentToByteArray(Document dom)
Serialize the Document object.- Parameters:
dom- the document to serialize- Returns:
- the serialized dom String
-
getTextValue
public static String getTextValue(Element element)
Returns the text value of an element (title, alt or contents). Note that the result is 50 characters or less in length.- Parameters:
element- The element.- Returns:
- The text value of the element.
-
getDifferences
public static List<org.custommonkey.xmlunit.Difference> getDifferences(String controlDom, String testDom)
Get differences between DOMs.- Parameters:
controlDom- The control dom.testDom- The test dom.- Returns:
- The differences.
-
getDifferences
public static List<org.custommonkey.xmlunit.Difference> getDifferences(String controlDom, String testDom, List<String> ignoreAttributes)
Get differences between DOMs.- Parameters:
controlDom- The control dom.testDom- The test dom.ignoreAttributes- The list of attributes to ignore.- Returns:
- The differences.
-
removeNewLines
public static String removeNewLines(String html)
Removes newlines from a string.- Parameters:
html- The string.- Returns:
- The new string without the newlines or tabs.
-
replaceString
public static String replaceString(String string, String regex, String replace)
- Parameters:
string- The original string.regex- The regular expression.replace- What to replace it with.- Returns:
- replaces regex in str by replace where the dot sign also supports newlines
-
addFolderSlashIfNeeded
public static String addFolderSlashIfNeeded(String folderName)
Adds a slash to a path if it doesn't end with a slash.- Parameters:
folderName- The path to append a possible slash.- Returns:
- The new, correct path.
-
getFileNameInPath
private static String getFileNameInPath(String path)
Returns the filename in a path. For example with path = "foo/bar/crawljax.txt" returns "crawljax.txt"- Parameters:
path-- Returns:
- the filename from the path
-
getTemplateAsString
public static String getTemplateAsString(String fileName) throws IOException
Retrieves the content of the filename. Also reads from JAR Searches for the resource in the root folder in the jar- Parameters:
fileName- Filename.- Returns:
- The contents of the file.
- Throws:
IOException- On error.
-
getFrameIdentification
public static String getFrameIdentification(Element frame)
- Parameters:
frame- the frame element.- Returns:
- the name or id of this element if they are present, otherwise null.
-
writeDocumentToFile
public static void writeDocumentToFile(Document document, String filePathname, String method, int indent) throws TransformerException, IOException
Write the document object to a file.- Parameters:
document- the document object.filePathname- the path name of the file to be written to.method- the output method: for instance html, xml, textindent- amount of indentation. -1 to use the default.- Throws:
TransformerException- if an exception occurs.IOException- if an IO exception occurs.
-
getAllLeafNodes
public static NodeList getAllLeafNodes(Document document) throws XPathExpressionException
- Throws:
XPathExpressionException
-
getAttributeFromElement
public static String getAttributeFromElement(Node element, String attribute)
-
getElementsByTagName
public static List<Node> getElementsByTagName(Document document, String tagName)
-
getTextContent
public static String getTextContent(Document document, boolean individualTokens)
To get all the textual content in the dom- Parameters:
document-individualTokens- : default True : when set to true, each text node from dom is used to build the text content : when set to false, the text content of whole is obtained at once.- Returns:
-
getDOMWithoutContent
public static String getDOMWithoutContent(Document document) throws XPathExpressionException
- Throws:
XPathExpressionException
-
getNumLeafNodes
public static int getNumLeafNodes(Node node) throws XPathExpressionException
- Throws:
XPathExpressionException
-
getAllSubtreeNodes
public static NodeList getAllSubtreeNodes(Node node) throws XPathExpressionException
- Throws:
XPathExpressionException
-
toUniformDOM
private static String toUniformDOM(String html)
- Parameters:
html- The html string.- Returns:
- uniform version of dom with predefined attributes stripped
-
filterAttributes
private static String filterAttributes(String html)
Filters attributes from the HTML string.- Parameters:
html- The HTML to filter.- Returns:
- The filtered HTML string.
-
contains
public static boolean contains(Node parent, Node child)
returns true if parent and child are the same node- Parameters:
parent-child-- Returns:
-
-