Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Each session duration is 2 hours.
Day-1: Session -1: Business Context for Big Data Business Intelligence in Government
- Case Studies from NIH and DoE
- Adoption rates of Big Data in government agencies and the alignment of future operations with Big Data Predictive Analytics
- Broad-scale application areas in DoD, NSA, IRS, USDA, and other entities
- Integration of Big Data with legacy data systems
- Foundational understanding of enabling technologies for predictive analytics
- Data integration and dashboard visualization
- Fraud management strategies
- Generation of business rules and fraud detection models
- Threat detection and profiling
- Cost-benefit analysis for Big Data implementation
Day-1: Session-2: Introduction to Big Data - Part 1
- Core characteristics of Big Data: volume, variety, velocity, and veracity. MPP architecture for handling volume.
- Data Warehouses – static schema with slowly evolving datasets
- MPP Databases such as Greenplum, Exadata, Teradata, Netezza, Vertica, and others
- Hadoop-Based Solutions – flexibility without rigid structure requirements for datasets
- Typical workflow pattern: HDFS, MapReduce (crunch), and retrieval from HDFS
- Batch processing – suited for analytical and non-interactive tasks
- Volume management via CEP streaming data
- Common choices – CEP products (e.g., Infostreams, Apama, MarkLogic, etc.)
- Less production-ready options – Storm/S4
- NoSQL Databases – (columnar and key-value): Ideal as an analytical adjunct to data warehouses/databases
Day-1: Session -3: Introduction to Big Data - Part 2
NoSQL Solutions
- KV Store - Keyspace, Flare, SchemaFree, RAMCloud, Oracle NoSQL Database (OnDB)
- KV Store - Dynamo, Voldemort, Dynomite, SubRecord, Mo8onDb, DovetailDB
- KV Store (Hierarchical) - GT.m, Cache
- KV Store (Ordered) - TokyoTyrant, Lightcloud, NMDB, Luxio, MemcacheDB, Actord
- KV Cache - Memcached, Repcached, Coherence, Infinispan, EXtremeScale, JBossCache, Velocity, Terracoqua
- Tuple Store - Gigaspaces, Coord, Apache River
- Object Database - ZopeDB, DB40, Shoal
- Document Store - CouchDB, Cloudant, Couchbase, MongoDB, Jackrabbit, XML-Databases, ThruDB, CloudKit, Prsevere, Riak-Basho, Scalaris
- Wide Columnar Store - BigTable, HBase, Apache Cassandra, Hypertable, KAI, OpenNeptune, Qbase, KDI
Data Varieties: Introduction to Data Cleaning Challenges in Big Data
- RDBMS – static structure/schema, which limits agile, exploratory environments
- NoSQL – semi-structured, offering sufficient structure to store data without a predefined exact schema
- Key data cleaning issues
Day-1: Session-4: Big Data Introduction - Part 3: Hadoop
- Criteria for selecting Hadoop
- STRUCTURED - Enterprise data warehouses/databases can store massive data (at a cost) but impose structure (limiting active exploration)
- SEMI-STRUCTURED data – difficult to manage with traditional solutions (DW/DB)
- Data warehousing requires significant effort and remains static even after implementation
- For handling data variety and volume, processed on commodity hardware – HADOOP
- Commodity hardware required to create a Hadoop Cluster
Introduction to MapReduce / HDFS
- MapReduce – distributing computing tasks over multiple servers
- HDFS – making data available locally for the computing process (with redundancy)
- Data – can be unstructured/schema-less (unlike RDBMS)
- Developer responsibility for interpreting data
- Programming MapReduce involves working with Java (pros/cons) and manually loading data into HDFS
Day-2: Session-1: The Big Data Ecosystem – Building Big Data ETL: The Universe of Big Data Tools – When to Use Which?
- Hadoop compared to other NoSQL solutions
- Use cases for interactive, random access to data
- Hbase (column-oriented database) built on top of Hadoop
- Random access capabilities with specific restrictions (max 1 PB)
- Limitations for ad-hoc analytics; suitability for logging, counting, and time-series data
- Sqoop – Importing from databases to Hive or HDFS (JDBC/ODBC access)
- Flume – Streaming data (e.g., log data) into HDFS
Day-2: Session-2: Big Data Management Systems
- Managing moving parts and node failures: ZooKeeper – for configuration, coordination, and naming services
- Managing complex pipelines/workflows: Oozie – for managing workflows, dependencies, and daisy-chaining tasks
- System administration tasks such as deployment, configuration, cluster management, and upgrades: Ambari
- Cloud deployment: Whirr
Day-2: Session-3: Predictive Analytics in Business Intelligence - Part 1: Fundamental Techniques & Machine Learning-based BI
- Introduction to Machine Learning
- Learning classification techniques
- Bayesian Prediction – preparing training files
- Support Vector Machines
- KNN p-Tree Algebra & vertical mining
- Neural Networks
- Addressing Big Data large variable problems – Random Forest (RF)
- Addressing Big Data automation problems – Multi-model ensemble RF
- Automation via Soft10-M
- Text analysis tool – Treeminer
- Agile learning
- Agent-based learning
- Distributed learning
- Introduction to open-source tools for predictive analytics: R, Rapidminer, Mahut
Day-2: Session-4: Predictive Analytics Ecosystem - Part 2: Common Predictive Analytics Challenges in Government
- Insight analytics
- Visualization analytics
- Structured predictive analytics
- Unstructured predictive analytics
- Threat/fraudster/vendor profiling
- Recommendation engines
- Pattern detection
- Rule/scenario discovery – failure, fraud, optimization
- Root cause discovery
- Sentiment analysis
- CRM analytics
- Network analytics
- Text analytics
- Technology-assisted review
- Fraud analytics
- Real-time analytics
Day-3: Session-1: Real-Time and Scalable Analytics on Hadoop
- Why common analytics algorithms fail in Hadoop/HDFS environments
- Apache Hama – for Bulk Synchronous distributed computing
- Apache Spark – for cluster computing in real-time analytics
- CMU Graphics Lab2 – Graph-based asynchronous approach to distributed computing
- KNN p-Algebra based approach from Treeminer for reducing operational hardware costs
Day-3: Session-2: Tools for eDiscovery and Forensics
- eDiscovery on Big Data vs. Legacy data – comparing cost and performance
- Predictive coding and technology-assisted review (TAR)
- Live demonstration of a TAR product (vMiner) to illustrate how TAR facilitates faster discovery
- Faster indexing through HDFS – addressing data velocity
- NLP or Natural Language Processing – various techniques and open-source products
- eDiscovery in foreign languages – technology for foreign language processing
Day-3: Session 3: Big Data BI for Cyber Security – Understanding the Full 360-Degree View from Rapid Data Collection to Threat Identification
- Foundations of security analytics – attack surface, security misconfiguration, host defenses
- Network infrastructure, large data pipes, and Response ETL for real-time analytics
- Prescriptive vs. predictive – Fixed rule-based systems vs. auto-discovery of threat rules from metadata
Day-3: Session 4: Big Data in USDA: Applications in Agriculture
- Introduction to IoT (Internet of Things) for agriculture – sensor-based Big Data and control
- Introduction to Satellite imaging and its applications in agriculture
- Integrating sensor and image data for soil fertility analysis, cultivation recommendations, and forecasting
- Agriculture insurance and Big Data
- Crop loss forecasting
Day-4: Session-1: Fraud Prevention BI from Big Data in Government – Fraud Analytics
- Basic classification of fraud analytics – rule-based vs. predictive analytics
- Supervised vs. unsupervised Machine Learning for fraud pattern detection
- Vendor fraud/overcharging for projects
- Medicare and Medicaid fraud – fraud detection techniques for claim processing
- Travel reimbursement fraud
- IRS refund fraud
- Case studies and live demos will be provided where data is available.
Day-4: Session-2: Social Media Analytics – Intelligence Gathering and Analysis
- Big Data ETL APIs for extracting social media data
- Handling text, image, metadata, and video
- Sentiment analysis from social media feeds
- Contextual and non-contextual filtering of social media feeds
- Social Media Dashboards for integrating diverse social media sources
- Automated profiling of social media profiles
- Live demonstrations of each analytics module using the Treeminer Tool.
Day-4: Session-3: Big Data Analytics in Image Processing and Video Feeds
- Image storage techniques in Big Data – storage solutions for data exceeding petabytes
- LTFS and LTO standards
- GPFS-LTFS (Layered storage solution for Big image data)
- Fundamentals of image analytics
- Object recognition
- Image segmentation
- Motion tracking
- 3-D image reconstruction
Day-4: Session-4: Big Data Applications in NIH
- Emerging areas in Bio-informatics
- Meta-genomics and Big Data mining challenges
- Big Data Predictive analytics for Pharmacogenomics, Metabolomics, and Proteomics
- Big Data in downstream Genomics processes
- Application of Big Data predictive analytics in Public health
Big Data Dashboards for Quick Accessibility of Diverse Data and Display
- Integrating existing application platforms with Big Data Dashboards
- Big Data management practices
- Case Study of Big Data Dashboards: Tableau and Pentaho
- Using Big Data apps to push location-based services in Government
- Tracking systems and management
Day-5: Session-1: Justifying Big Data BI Implementation Within an Organization
- Defining ROI for Big Data implementation
- Case studies on saving analyst time for data collection and preparation – productivity gains
- Case studies on revenue gains from reducing licensed database costs
- Revenue gains from location-based services
- Savings from fraud prevention
- An integrated spreadsheet approach to calculate approximate expense vs. revenue gain/savings from Big Data implementation.
Day-5: Session-2: Step-by-Step Procedure to Replace Legacy Data Systems with Big Data Systems
- Understanding a practical Big Data Migration Roadmap
- Key information required before architecting a Big Data implementation
- Methods for calculating volume, velocity, variety, and veracity of data
- Estimating data growth
- Case studies
Day-5: Session 4: Review of Big Data Vendors and Their Products. Q/A Session
- Accenture
- APTEAN (Formerly CDC Software)
- Cisco Systems
- Cloudera
- Dell
- EMC
- GoodData Corporation
- Guavus
- Hitachi Data Systems
- Hortonworks
- HP
- IBM
- Informatica
- Intel
- Jaspersoft
- Microsoft
- MongoDB (Formerly 10Gen)
- MU Sigma
- Netapp
- Opera Solutions
- Oracle
- Pentaho
- Platfora
- Qliktech
- Quantum
- Rackspace
- Revolution Analytics
- Salesforce
- SAP
- SAS Institute
- Sisense
- Software AG/Terracotta
- Soft10 Automation
- Splunk
- Sqrrl
- Supermicro
- Tableau Software
- Teradata
- Think Big Analytics
- Tidemark Systems
- Treeminer
- VMware (Part of EMC)
Requirements
- Fundamental knowledge of business operations and data systems within the relevant government domain
- Basic comprehension of SQL/Oracle or relational database concepts
- Foundational understanding of Statistics (at a spreadsheet level)
35 Hours
Testimonials (1)
The ability of the trainer to align the course with the requirements of the organization other than just providing the course for the sake of delivering it.