[vc_empty_space][vc_empty_space]
Scalable sequential pattern mining based on PrefixSpan for high dimensional data
Akbar M.N.a, Saptawati G.A.P.a
a School of Electronic Engineering and Informatics, Institute of Technology Bandung, Bandung, Indonesia
[vc_row][vc_column][vc_row_inner][vc_column_inner][vc_separator css=”.vc_custom_1624529070653{padding-top: 30px !important;padding-bottom: 30px !important;}”][/vc_column_inner][/vc_row_inner][vc_row_inner layout=”boxed”][vc_column_inner width=”3/4″ css=”.vc_custom_1624695412187{border-right-width: 1px !important;border-right-color: #dddddd !important;border-right-style: solid !important;border-radius: 1px !important;}”][vc_empty_space][megatron_heading title=”Abstract” size=”size-sm” text_align=”text-left”][vc_column_text]© 2016 IEEE.The phenomenon of data explosion makes analysis to find insights inside the data become more difficult. A problem that often occurs in the application of pattern recognition in the real-world domain is not only caused by the large size of data but also the high-dimensional data. Data analysis demands that a large and complex data can be processed quickly and optimally to support decision making. This study offered a scalable sequential patterns extraction to gain more insight from the data using PrefixSpan implemented on the Spark platform as a distributed system. The goal is to overcome the problem of increasing the amount of data (scalability) in complex and high dimensional data effectively and in a relatively quick performance. The experiments show that this method can make full use of cluster computing resources to accelerate the mining process, reduces the time of scanning database and build projected database with an increasing number of worker on the Spark platform.[/vc_column_text][vc_empty_space][vc_separator css=”.vc_custom_1624528584150{padding-top: 25px !important;padding-bottom: 25px !important;}”][vc_empty_space][megatron_heading title=”Author keywords” size=”size-sm” text_align=”text-left”][vc_column_text]Computing resource,Distributed systems,High dimensional data,Prefix spans,Projected database,Sequential patterns,Sequential patterns mining,Sequential-pattern mining[/vc_column_text][vc_empty_space][vc_separator css=”.vc_custom_1624528584150{padding-top: 25px !important;padding-bottom: 25px !important;}”][vc_empty_space][megatron_heading title=”Indexed keywords” size=”size-sm” text_align=”text-left”][vc_column_text]High dimensional data,PrefixSpan,Scalability,Sequential patterns mining,Spark[/vc_column_text][vc_empty_space][vc_separator css=”.vc_custom_1624528584150{padding-top: 25px !important;padding-bottom: 25px !important;}”][vc_empty_space][megatron_heading title=”Funding details” size=”size-sm” text_align=”text-left”][vc_column_text][/vc_column_text][vc_empty_space][vc_separator css=”.vc_custom_1624528584150{padding-top: 25px !important;padding-bottom: 25px !important;}”][vc_empty_space][megatron_heading title=”DOI” size=”size-sm” text_align=”text-left”][vc_column_text]https://doi.org/10.1109/ICODSE.2016.7936122[/vc_column_text][/vc_column_inner][vc_column_inner width=”1/4″][vc_column_text]Widget Plumx[/vc_column_text][/vc_column_inner][/vc_row_inner][/vc_column][/vc_row][vc_row][vc_column][vc_separator css=”.vc_custom_1624528584150{padding-top: 25px !important;padding-bottom: 25px !important;}”][/vc_column][/vc_row]