Skip to content
SmartStudy

Big Data, Cloud Computing and Data Architecture

Technology

Big Data, Cloud Computing and Data Architecture

Syllabus tag: KASNEB CPA | Advanced Level | CA34S1 Business Data Analytics

1. Defining big data — the Vs

Big data is characterised by the Three Vs (Gartner, 2001):

  • Volume: data volumes too large for traditional databases — terabytes, petabytes, or more
  • Velocity: data generated and must be processed at high speed (real-time or near-real-time streams)
  • Variety: diverse formats — structured (databases), semi-structured (JSON, XML), unstructured (text, images, video)

Extended models add: Veracity (trustworthiness and quality of the data) and Value (the business insight extracted).

2. Big data technologies

Hadoop: an open-source framework for distributed storage (HDFS — Hadoop Distributed File System) and processing (MapReduce) of large datasets across clusters of commodity hardware. Suited for batch processing.

Apache Spark: faster than Hadoop for many workloads — performs in-memory processing. Supports batch, streaming, machine learning (MLlib), and graph processing. Now the dominant big data processing framework.

Kafka: a distributed event streaming platform for real-time data pipelines — ingests high-velocity data (e.g. IoT sensor data, financial transactions) and makes it available to downstream systems.

NoSQL databases: designed for unstructured or semi-structured big data: document stores (MongoDB), key-value stores (Redis), column-family stores (Cassandra), graph databases (Neo4j).

3. Data architecture patterns

Data warehouse: a structured, integrated repository of historical data from multiple sources, optimised for analytical queries. Schema-on-write — data is structured before loading. Examples: Amazon Redshift, Snowflake, Google BigQuery.

Data lake: a centralised repository storing raw data in its native format until needed for analysis. Schema-on-read — structure is applied when the data is queried. More flexible; lower upfront cost; but risk of becoming a "data swamp" without governance.

Data lakehouse: combines the flexibility of a data lake with the structure and performance of a data warehouse. Modern architecture supported by platforms like Databricks and Delta Lake.

ETL (Extract, Transform, Load): the traditional process of moving data from source systems to a data warehouse. ELT (Extract, Load, Transform): modern cloud-native approach — load raw data into the data warehouse first, then transform in-place using the warehouse's compute power.

4. Cloud computing

Cloud service models: IaaS (Infrastructure as a Service — virtual machines, storage); PaaS (Platform as a Service — managed databases, development platforms); SaaS (Software as a Service — ready-to-use applications like Salesforce, M365).

Cloud advantages for analytics: scalability on demand; reduced upfront capital expenditure; access to managed analytics services; global availability; built-in redundancy and disaster recovery.

Major cloud providers: AWS (S3, Redshift, Athena, SageMaker); Microsoft Azure (Data Lake, Synapse, ML Studio); Google Cloud (BigQuery, Vertex AI, Looker).

5. Real-time vs batch processing

Batch processing: data is collected over a period and processed in bulk at scheduled intervals (e.g. nightly batch jobs). Efficient for large volumes; results are not immediate. Stream processing: data is processed continuously as it arrives (e.g. fraud detection on card transactions, real-time dashboards). More complex to implement; enables immediate insights and actions.

Next in Business Data AnalyticsMachine Learning and Artificial Intelligence →