A Text Extraction Method Based on Keyword Matching
A keyword matching and keyword technology, applied to instruments, other database retrieval, calculation, etc., can solve the problems of difficult web page text extraction, high error rate, strong dependence on website structure, etc., to ensure objectivity and rationality, Good versatility, simplified traversal effect
- Summary
- Abstract
- Description
- Claims
- Application Information
AI Technical Summary
Problems solved by technology
Method used
Image
Examples
Embodiment Construction
[0041] The content of the present invention will be further elaborated below in conjunction with the accompanying drawings, but it is not intended to limit the present invention.
[0042] Such as figure 1 As shown, the text extraction method based on keyword matching of the present invention specifically includes the following steps:
[0043] (1) Webpage preprocessing, counting and extracting the keywords in the Keywords tag of the webpage source code, and establishing a standard library with keywords; using regular expressions to preprocess the webpage to be processed, removing obvious noise text, and obtaining a rough webpage;
[0044] (2) Build a DOM tree, use the Jsoup tool to parse the HTML of the rough web page, and obtain the data of the rough web page; DOM uses a set of structured nodes and objects to represent the structure of the document, that is, each component in the document is defined as a node , so as to connect the webpage, scripting language and pr...
PUM
Abstract
Description
Claims
Application Information
- R&D Engineer
- R&D Manager
- IP Professional
- Industry Leading Data Capabilities
- Powerful AI technology
- Patent DNA Extraction
Browse by: Latest US Patents, China's latest patents, Technical Efficacy Thesaurus, Application Domain, Technology Topic, Popular Technical Reports.
© 2024 PatSnap. All rights reserved.Legal|Privacy policy|Modern Slavery Act Transparency Statement|Sitemap|About US| Contact US: help@patsnap.com