Get in Touch

Course Outline

Each session duration is 2 hours.

Day-1: Session -1: Business Context for Big Data Business Intelligence in Government

  • Case Studies from NIH and DoE
  • Adoption rates of Big Data in government agencies and the alignment of future operations with Big Data Predictive Analytics
  • Broad-scale application areas in DoD, NSA, IRS, USDA, and other entities
  • Integration of Big Data with legacy data systems
  • Foundational understanding of enabling technologies for predictive analytics
  • Data integration and dashboard visualization
  • Fraud management strategies
  • Generation of business rules and fraud detection models
  • Threat detection and profiling
  • Cost-benefit analysis for Big Data implementation

Day-1: Session-2: Introduction to Big Data - Part 1

  • Core characteristics of Big Data: volume, variety, velocity, and veracity. MPP architecture for handling volume.
  • Data Warehouses – static schema with slowly evolving datasets
  • MPP Databases such as Greenplum, Exadata, Teradata, Netezza, Vertica, and others
  • Hadoop-Based Solutions – flexibility without rigid structure requirements for datasets
  • Typical workflow pattern: HDFS, MapReduce (crunch), and retrieval from HDFS
  • Batch processing – suited for analytical and non-interactive tasks
  • Volume management via CEP streaming data
  • Common choices – CEP products (e.g., Infostreams, Apama, MarkLogic, etc.)
  • Less production-ready options – Storm/S4
  • NoSQL Databases – (columnar and key-value): Ideal as an analytical adjunct to data warehouses/databases

Day-1: Session -3: Introduction to Big Data - Part 2

NoSQL Solutions

  • KV Store - Keyspace, Flare, SchemaFree, RAMCloud, Oracle NoSQL Database (OnDB)
  • KV Store - Dynamo, Voldemort, Dynomite, SubRecord, Mo8onDb, DovetailDB
  • KV Store (Hierarchical) - GT.m, Cache
  • KV Store (Ordered) - TokyoTyrant, Lightcloud, NMDB, Luxio, MemcacheDB, Actord
  • KV Cache - Memcached, Repcached, Coherence, Infinispan, EXtremeScale, JBossCache, Velocity, Terracoqua
  • Tuple Store - Gigaspaces, Coord, Apache River
  • Object Database - ZopeDB, DB40, Shoal
  • Document Store - CouchDB, Cloudant, Couchbase, MongoDB, Jackrabbit, XML-Databases, ThruDB, CloudKit, Prsevere, Riak-Basho, Scalaris
  • Wide Columnar Store - BigTable, HBase, Apache Cassandra, Hypertable, KAI, OpenNeptune, Qbase, KDI

Data Varieties: Introduction to Data Cleaning Challenges in Big Data

  • RDBMS – static structure/schema, which limits agile, exploratory environments
  • NoSQL – semi-structured, offering sufficient structure to store data without a predefined exact schema
  • Key data cleaning issues

Day-1: Session-4: Big Data Introduction - Part 3: Hadoop

  • Criteria for selecting Hadoop
  • STRUCTURED - Enterprise data warehouses/databases can store massive data (at a cost) but impose structure (limiting active exploration)
  • SEMI-STRUCTURED data – difficult to manage with traditional solutions (DW/DB)
  • Data warehousing requires significant effort and remains static even after implementation
  • For handling data variety and volume, processed on commodity hardware – HADOOP
  • Commodity hardware required to create a Hadoop Cluster

Introduction to MapReduce / HDFS

  • MapReduce – distributing computing tasks over multiple servers
  • HDFS – making data available locally for the computing process (with redundancy)
  • Data – can be unstructured/schema-less (unlike RDBMS)
  • Developer responsibility for interpreting data
  • Programming MapReduce involves working with Java (pros/cons) and manually loading data into HDFS

Day-2: Session-1: The Big Data Ecosystem – Building Big Data ETL: The Universe of Big Data Tools – When to Use Which?

  • Hadoop compared to other NoSQL solutions
  • Use cases for interactive, random access to data
  • Hbase (column-oriented database) built on top of Hadoop
  • Random access capabilities with specific restrictions (max 1 PB)
  • Limitations for ad-hoc analytics; suitability for logging, counting, and time-series data
  • Sqoop – Importing from databases to Hive or HDFS (JDBC/ODBC access)
  • Flume – Streaming data (e.g., log data) into HDFS

Day-2: Session-2: Big Data Management Systems

  • Managing moving parts and node failures: ZooKeeper – for configuration, coordination, and naming services
  • Managing complex pipelines/workflows: Oozie – for managing workflows, dependencies, and daisy-chaining tasks
  • System administration tasks such as deployment, configuration, cluster management, and upgrades: Ambari
  • Cloud deployment: Whirr

Day-2: Session-3: Predictive Analytics in Business Intelligence - Part 1: Fundamental Techniques & Machine Learning-based BI

  • Introduction to Machine Learning
  • Learning classification techniques
  • Bayesian Prediction – preparing training files
  • Support Vector Machines
  • KNN p-Tree Algebra & vertical mining
  • Neural Networks
  • Addressing Big Data large variable problems – Random Forest (RF)
  • Addressing Big Data automation problems – Multi-model ensemble RF
  • Automation via Soft10-M
  • Text analysis tool – Treeminer
  • Agile learning
  • Agent-based learning
  • Distributed learning
  • Introduction to open-source tools for predictive analytics: R, Rapidminer, Mahut

Day-2: Session-4: Predictive Analytics Ecosystem - Part 2: Common Predictive Analytics Challenges in Government

  • Insight analytics
  • Visualization analytics
  • Structured predictive analytics
  • Unstructured predictive analytics
  • Threat/fraudster/vendor profiling
  • Recommendation engines
  • Pattern detection
  • Rule/scenario discovery – failure, fraud, optimization
  • Root cause discovery
  • Sentiment analysis
  • CRM analytics
  • Network analytics
  • Text analytics
  • Technology-assisted review
  • Fraud analytics
  • Real-time analytics

Day-3: Session-1: Real-Time and Scalable Analytics on Hadoop

  • Why common analytics algorithms fail in Hadoop/HDFS environments
  • Apache Hama – for Bulk Synchronous distributed computing
  • Apache Spark – for cluster computing in real-time analytics
  • CMU Graphics Lab2 – Graph-based asynchronous approach to distributed computing
  • KNN p-Algebra based approach from Treeminer for reducing operational hardware costs

Day-3: Session-2: Tools for eDiscovery and Forensics

  • eDiscovery on Big Data vs. Legacy data – comparing cost and performance
  • Predictive coding and technology-assisted review (TAR)
  • Live demonstration of a TAR product (vMiner) to illustrate how TAR facilitates faster discovery
  • Faster indexing through HDFS – addressing data velocity
  • NLP or Natural Language Processing – various techniques and open-source products
  • eDiscovery in foreign languages – technology for foreign language processing

Day-3: Session 3: Big Data BI for Cyber Security – Understanding the Full 360-Degree View from Rapid Data Collection to Threat Identification

  • Foundations of security analytics – attack surface, security misconfiguration, host defenses
  • Network infrastructure, large data pipes, and Response ETL for real-time analytics
  • Prescriptive vs. predictive – Fixed rule-based systems vs. auto-discovery of threat rules from metadata

Day-3: Session 4: Big Data in USDA: Applications in Agriculture

  • Introduction to IoT (Internet of Things) for agriculture – sensor-based Big Data and control
  • Introduction to Satellite imaging and its applications in agriculture
  • Integrating sensor and image data for soil fertility analysis, cultivation recommendations, and forecasting
  • Agriculture insurance and Big Data
  • Crop loss forecasting

Day-4: Session-1: Fraud Prevention BI from Big Data in Government – Fraud Analytics

  • Basic classification of fraud analytics – rule-based vs. predictive analytics
  • Supervised vs. unsupervised Machine Learning for fraud pattern detection
  • Vendor fraud/overcharging for projects
  • Medicare and Medicaid fraud – fraud detection techniques for claim processing
  • Travel reimbursement fraud
  • IRS refund fraud
  • Case studies and live demos will be provided where data is available.

Day-4: Session-2: Social Media Analytics – Intelligence Gathering and Analysis

  • Big Data ETL APIs for extracting social media data
  • Handling text, image, metadata, and video
  • Sentiment analysis from social media feeds
  • Contextual and non-contextual filtering of social media feeds
  • Social Media Dashboards for integrating diverse social media sources
  • Automated profiling of social media profiles
  • Live demonstrations of each analytics module using the Treeminer Tool.

Day-4: Session-3: Big Data Analytics in Image Processing and Video Feeds

  • Image storage techniques in Big Data – storage solutions for data exceeding petabytes
  • LTFS and LTO standards
  • GPFS-LTFS (Layered storage solution for Big image data)
  • Fundamentals of image analytics
  • Object recognition
  • Image segmentation
  • Motion tracking
  • 3-D image reconstruction

Day-4: Session-4: Big Data Applications in NIH

  • Emerging areas in Bio-informatics
  • Meta-genomics and Big Data mining challenges
  • Big Data Predictive analytics for Pharmacogenomics, Metabolomics, and Proteomics
  • Big Data in downstream Genomics processes
  • Application of Big Data predictive analytics in Public health

Big Data Dashboards for Quick Accessibility of Diverse Data and Display

  • Integrating existing application platforms with Big Data Dashboards
  • Big Data management practices
  • Case Study of Big Data Dashboards: Tableau and Pentaho
  • Using Big Data apps to push location-based services in Government
  • Tracking systems and management

Day-5: Session-1: Justifying Big Data BI Implementation Within an Organization

  • Defining ROI for Big Data implementation
  • Case studies on saving analyst time for data collection and preparation – productivity gains
  • Case studies on revenue gains from reducing licensed database costs
  • Revenue gains from location-based services
  • Savings from fraud prevention
  • An integrated spreadsheet approach to calculate approximate expense vs. revenue gain/savings from Big Data implementation.

Day-5: Session-2: Step-by-Step Procedure to Replace Legacy Data Systems with Big Data Systems

  • Understanding a practical Big Data Migration Roadmap
  • Key information required before architecting a Big Data implementation
  • Methods for calculating volume, velocity, variety, and veracity of data
  • Estimating data growth
  • Case studies

Day-5: Session 4: Review of Big Data Vendors and Their Products. Q/A Session

  • Accenture
  • APTEAN (Formerly CDC Software)
  • Cisco Systems
  • Cloudera
  • Dell
  • EMC
  • GoodData Corporation
  • Guavus
  • Hitachi Data Systems
  • Hortonworks
  • HP
  • IBM
  • Informatica
  • Intel
  • Jaspersoft
  • Microsoft
  • MongoDB (Formerly 10Gen)
  • MU Sigma
  • Netapp
  • Opera Solutions
  • Oracle
  • Pentaho
  • Platfora
  • Qliktech
  • Quantum
  • Rackspace
  • Revolution Analytics
  • Salesforce
  • SAP
  • SAS Institute
  • Sisense
  • Software AG/Terracotta
  • Soft10 Automation
  • Splunk
  • Sqrrl
  • Supermicro
  • Tableau Software
  • Teradata
  • Think Big Analytics
  • Tidemark Systems
  • Treeminer
  • VMware (Part of EMC)

Requirements

  • Fundamental knowledge of business operations and data systems within the relevant government domain
  • Basic comprehension of SQL/Oracle or relational database concepts
  • Foundational understanding of Statistics (at a spreadsheet level)
 35 Hours

Number of participants


Price per participant

Testimonials (1)

Upcoming Courses

Related Categories