Lesson 4.11.1.1

4.11.1.1 Big Data Quiz: AQA Computer Science, Unit 11

20 questions

In partnership with Revision Ninja

Lesson 4.11.1.1, Big Data: 20 multiple choice questions for the AQA Computer Science (7517), Unit 11: Big Data, written with Revision Ninja.

Host it live on the board and students join with a game code on their own devices, or revise alone with Free Play. The answers are revealed in the game.

Host this setFree Play

The 20 questions

  1. Which of these is one of the three characteristics of Big Data described in the specification?

    • Virtualisation
    • Velocity
    • Vendor
    • Validation
  2. What does volume mean when describing Big Data?

    • Data that is too big to fit into a single server.
    • The number of different file types that are stored in one database.
    • The number of users who connect to a web server each day.
    • The loudness of a data stream that is captured by a sensor in the field.
  3. What does velocity mean when describing Big Data?

    • The number of fields contained in each record of a relational table.
    • Streaming data that must be responded to within milliseconds to seconds.
    • The size in bytes of each individual file uploaded to a server.
    • The speed at which a processor completes one single machine instruction, measured in cycles per second on the processor's clock.
  4. What does variety mean when describing Big Data?

    • Data that is always stored as identical rows and columns in a single table.
    • Data in many forms, such as structured, unstructured, text and multimedia.
    • Many different database vendors offering products that are very similar to each other.
    • The number of different programming languages that a development team is able to use.
  5. Why are relational databases often not appropriate for much Big Data?

    • They are only available for single-user systems and cannot be shared between users, so each person needs a separate copy of the data.
    • They require data to fit into rows and columns and do not scale well across many machines.
    • They are always too slow to read any data, however small the data set may be.
    • They cannot store any numeric values, and Big Data is made up mostly of numbers.
  6. Which technique is needed to find patterns in large unstructured data sets?

    • Machine learning techniques.
    • Client-side rendering of the data inside a thin client device.
    • Manual data entry by a team of clerks checking each record one by one.
    • Normalisation of the data to third normal form in a relational schema.
  7. When data is too big to fit on a single server, how must the processing be handled?

    • It must be carried out by the printer queue of the server, one job at a time.
    • It must be limited to a single core of one processor in the server room.
    • It must be stored in one file of unlimited size on the server's local disk.
    • It must be distributed across more than one machine.
  8. Which features of functional programming make it easier to write correct code?

    • Long sequences of jump statements that move control between sections of code.
    • Immutable data structures and statelessness.
    • Mutable shared state that many threads update at the same time.
    • Global variables that every function is free to change at any moment.
  9. Which feature of functional programming supports running code on many servers?

    • Mutable objects that each server changes in place as it processes the data.
    • Statelessness, so functions can run on any machine without relying on shared state.
    • Reliance on a shared global variable that every server must update in turn.
    • Procedures that depend on the exact order of statements to produce each result, so that every statement must run in a fixed sequence.
  10. In a fact-based model, what does each fact capture?

    • An entire table of related records stored together in a single row.
    • A complete copy of the database schema for every user of the system.
    • A single piece of information.
    • The name of the programming language that is used to query the dataset.
  11. In a graph schema, what do nodes represent?

    • The physical servers on which the data set is stored.
    • The connections between two records held in a relational table.
    • Entities in the data set, such as a person or a product.
    • The data types assigned to each attribute of the schema.
  12. In a graph schema, what does an edge represent?

    • A duplicate copy of a node that is kept for backup purposes only.
    • A property that describes a single node in the data set in detail.
    • The identifier of the server on which the node is currently stored.
    • A relationship between two nodes.
  13. In a graph schema, what is a property?

    • A link that always connects one node to exactly one other node in the graph.
    • An attribute value attached to a node or an edge, such as a name or a date.
    • A table that stores all of the nodes of the graph in their original order.
    • A rule that removes duplicate edges from the data set whenever it is loaded.
  14. In a social network graph, Alice follows Bob since 2020. Which best describes this?

    • A property on Alice's node that stores the whole record for Bob.
    • A node labelled 2020 that is connected to both Alice and Bob in the graph.
    • Two separate nodes, one for Alice and one for Bob, with no link between them, and each node stores the date of its creation as a property.
    • An edge between the Alice and Bob nodes, with a property recording the year 2020.
  15. A dataset holds sensor readings streamed to the server every second. Which Big Data characteristic is most prominent?

    • Relational integrity, because each reading has a foreign key pointing to a sensor.
    • Velocity
    • Variety, because the readings always come in one fixed format from every sensor.
    • Normalisation, because the readings must be split into several related tables.
  16. A dataset mixes video, text comments and structured sales records. Which characteristic is most evident?

    • Velocity, because the sales records are updated every second on the server.
    • Variety
    • Volume, because the video files are small in size compared with the records.
    • Primary key uniqueness, because every record in the data set has an identifier.
  17. A company holds a data set too large for one server. Which approach is appropriate?

    • Distribute the data and processing across several machines, using functional programming to keep the code correct.
    • Remove the unstructured data so that the remaining data fits on one server.
    • Upgrade the processor of the single server and keep all of the data in one table.
    • Store all of the data in one spreadsheet file on a local hard drive of a desktop.
  18. Why do immutable data structures help when processing is distributed?

    • They mean that each machine must lock the entire data set before it reads any value.
    • They allow every machine to edit the same value at the same time without any coordination.
    • They ensure that all data is always stored in a relational table format on each machine.
    • Values never change once created, so copies held on different machines stay consistent.
  19. Evaluate: why are unstructured data sources hard to analyse with relational databases?

    • They must be stored in a graph schema, which relational systems do not support at all.
    • Their content does not fit into fixed rows and columns, so patterns must be found by other means such as machine learning.
    • They are always stored in binary form, which relational systems are unable to read.
    • They have too few attributes to be meaningful when placed in a table structure.
  20. A stateless function given the same input always returns the same output. Why does this help in distributed systems?

    • The function can only run on the server that originally created the data it uses.
    • Any server can run the function without needing to know the state held on other servers.
    • It removes the need for any networking between the servers in the cluster at all.
    • It guarantees that each server stores a copy of every previous result it has produced, so that results never need to be recalculated.

All AQA Computer Science quizzes