Data-Mining
Data-Mining is the computational process of discovering patterns in large data sets involving methods at the intersection of Machine Learning, Statistics, and Database Systems. It is an essential step in the Knowledge Discovery in Databases (KDD) process, which aims to extract information from a data set and transform it into an understandable structure for further use. According to Wikipedia, the term is a misnomer, as the goal is the extraction of patterns and knowledge from large amounts of data, not the extraction of data itself.
Core Methodologies
The field utilizes several key techniques to process information. Clustering is used to find groups and structures in the data that are in some way similar, without using known structures. Classification involves generalizing known structures to apply to new data, such as identifying an email as spam. Regression-Analysis seeks to find a function which models the data with the least error. Industry leaders like IBM emphasize that these techniques allow businesses to reduce risks and increase revenue.
Applications and Tools
Modern practitioners often leverage Artificial Intelligence to automate complex discovery tasks. Popular programming environments for these tasks include Python and R-Programming, which offer extensive libraries for Predictive-Analytics. For large-scale enterprise operations, Oracle provides integrated solutions that combine mining capabilities with Data-Warehousing.