[vc_empty_space][vc_empty_space]
Information extraction on novel text using machine learning and rule-based system
Chaniago R.a, Khodra M.L.a
a School of Electrical Engineering and Informatics, Bandung Institute of Technology, Bandung, Indonesia
[vc_row][vc_column][vc_row_inner][vc_column_inner][vc_separator css=”.vc_custom_1624529070653{padding-top: 30px !important;padding-bottom: 30px !important;}”][/vc_column_inner][/vc_row_inner][vc_row_inner layout=”boxed”][vc_column_inner width=”3/4″ css=”.vc_custom_1624695412187{border-right-width: 1px !important;border-right-color: #dddddd !important;border-right-style: solid !important;border-radius: 1px !important;}”][vc_empty_space][megatron_heading title=”Abstract” size=”size-sm” text_align=”text-left”][vc_column_text]© 2017 IEEE.Novel consists of around 30,000 to 50,000 words in total. It usually tells a story about entities and its relation one another such as, Person, Location or Organization. In order to apprehend those information, reading the whole novel is compulsory. However, it is a time-consuming task. This research proposes a solution – automatic extraction of entity relation by means of Information Extraction (IE) technique. This technique is divided into two steps. First, all the entities are retrieved from the text input, by using Named Entity Recognition (NER). Afterward, all relations is extracted by Relation Extraction (RE) process. This research implements an IE system to both NER and RE, which employs supervised machine learning approach combined with rule-based system. The main purpose of this research is to determine which features and algorithm of the machine learning are adequate to acquire the best result, and which rules are the most suitable for novel characteristics.[/vc_column_text][vc_empty_space][vc_separator css=”.vc_custom_1624528584150{padding-top: 25px !important;padding-bottom: 25px !important;}”][vc_empty_space][megatron_heading title=”Author keywords” size=”size-sm” text_align=”text-left”][vc_column_text]Automatic extraction,Information extraction techniques,Named entity recognition,Relation extraction,Supervised machine learning,Text input,Time-consuming tasks[/vc_column_text][vc_empty_space][vc_separator css=”.vc_custom_1624528584150{padding-top: 25px !important;padding-bottom: 25px !important;}”][vc_empty_space][megatron_heading title=”Indexed keywords” size=”size-sm” text_align=”text-left”][vc_column_text]Information Extraction,Machine Learning,Named Entity Recognition,Relation Extraction[/vc_column_text][vc_empty_space][vc_separator css=”.vc_custom_1624528584150{padding-top: 25px !important;padding-bottom: 25px !important;}”][vc_empty_space][megatron_heading title=”Funding details” size=”size-sm” text_align=”text-left”][vc_column_text][/vc_column_text][vc_empty_space][vc_separator css=”.vc_custom_1624528584150{padding-top: 25px !important;padding-bottom: 25px !important;}”][vc_empty_space][megatron_heading title=”DOI” size=”size-sm” text_align=”text-left”][vc_column_text]https://doi.org/10.1109/INNOCIT.2017.8319148[/vc_column_text][/vc_column_inner][vc_column_inner width=”1/4″][vc_column_text]Widget Plumx[/vc_column_text][/vc_column_inner][/vc_row_inner][/vc_column][/vc_row][vc_row][vc_column][vc_separator css=”.vc_custom_1624528584150{padding-top: 25px !important;padding-bottom: 25px !important;}”][/vc_column][/vc_row]