The state of most data analytics pipelines is deplorable: too little automation; minimal reuse of code and data; and lack of coordination. The result is poor quality data delivered too late to meet business needs. DataOps is an emerging approach for building data pipelines and solutions. In this session Wayne Eckerson will explore trends in DataOps adoption, challenges and best practices.
It is notoriously difficult to achieve the promise of self-service analytics. Wayne Eckerson will explain how to empower business users to create their own reports and models without creating data chaos. He will show how to build a self-sustaining analytical culture that balances speed and standards, agility and architecture, and self-service and governance.
Data Governance is one of the hottest topics in data management, focusing both on how Governance driven change can enable companies to gain better leverage from their data and also to help them design and enforce the controls needed to ensure they remain compliant with regulations. Despite this rapidly growing focus, many Data Governance initiatives fail to meet their goals. This session will outline why Data Governance and Architecture should be connected, how to make it happen, and what part Business Intelligence and Data Warehousing will play in defining a robust and sustainable Governance programme.
In Data preparation it is important how the human modeler creates a dataset that is uniquely suited to the business problem. In this session Keith McCormick will expose analytic practitioners, data scientists, and those looking to get started in predictive analytics to the critical importance of properly preparing data in advance of model building. He will present the critical role of feature engineering and explaining how to do it effectively.
ASML, manufacturer of machines for the production of semiconductors, is implementing a central data lake to capture this data and make it accessible for reporting and analytics in a central environment. The data lake environment also includes an analytics lab for detailed exploration of data. In this session Jeroen Vermunt presents real-life examples of how ASML approaches the challenges of managing rapidly changing data.
Ensembling is one of the hottest techniques in today’s predictive analytics competitions. Every single recent winner of Kaggle.com and KDD competitions used an ensemble technique, including famous algorithms such as XGBoost and Random Forest. This session will provide a detailed overview of ensemble models, their origin, and show why they are so effective.
Migrating an existing data warehouse to the cloud is a complex process of moving schema, data, and ETL. The complexity increases when architectural modernization, restructuring of database schema or rebuilding of data pipelines is needed. In this session Dave Wells provides an overview of the benefits, techniques, and challenges when migrating an existing data warehouse to the cloud.