PageRank and information entropy-based text word segmentation method for judgment document
A text word segmentation and information entropy technology, applied in special data processing applications, instruments, electrical digital data processing, etc., can solve problems such as difficult recognition
- Summary
- Abstract
- Description
- Claims
- Application Information
AI Technical Summary
Problems solved by technology
Method used
Image
Examples
Embodiment Construction
[0058] The present invention mainly uses the improved PageRank algorithm to establish a graphical model of the inclusion relationship between latent words, and calculates the Rank value of all latent words and combines information entropy and mutual information to carry out word segmentation. The present invention adds a keyword dictionary to improve Adapt well to terminology in different domains. The overall process of the word segmentation method is as follows figure 1 shown. Its specific implementation steps are as follows:
[0059] 1. The main flow of the method is as follows: Figure 10 shown in the upper part.
[0060] Step (1), read the input text, segment it with punctuation marks, numbers and English letters as separators to get all the Chinese characters in the text, and then filter and remove the words with a word length of only 1 to get a string list S;
[0061] Step (2), for each string S in S i A substring S whose length does not exceed k (k=6) sub (potenti...
PUM
Abstract
Description
Claims
Application Information
- R&D Engineer
- R&D Manager
- IP Professional
- Industry Leading Data Capabilities
- Powerful AI technology
- Patent DNA Extraction
Browse by: Latest US Patents, China's latest patents, Technical Efficacy Thesaurus, Application Domain, Technology Topic, Popular Technical Reports.
© 2024 PatSnap. All rights reserved.Legal|Privacy policy|Modern Slavery Act Transparency Statement|Sitemap|About US| Contact US: help@patsnap.com