Dear learning community,
as many of you asked for some more details on the correct answers for the closed assignments or access to the past questions if you missed them, we decided to put the solutions online in the respective week.
If there are any questions left, please do not hesitate to ask in the forum.
Best regards,
Ralf for the IMDB Teaching Team
Question Name: Data Persistence
Text: Which data medium provides the best requirements to guarantee the persistence of the data?
Correct answer:
- SSD
Incorrect answers:
- DRAM
- CPU
- GPU
General explanation: Solid State Drives (SSD) offer non-volatile storage, while the others lose stored data without a steady power connection. In general, CPUs (Central Processing Units) and GPUs (Graphics Processing Units) are not storage but processing devices, as the name already indicates.
Question Name: Logging Method
Text: Which of these choices is a logging method?
Correct answer:
- Logical Logging
Incorrect answers:
- Reversal Logging
- Precipitate Logging
- Common Logging
General explanation: The simplest way of logging is to write the SQL statement and its parameters, the record id and the attribute values, to disk. This is called Logical Logging. The three other answers are made up.
Question Name: Recovery Process
Text: When recovering an in-memory database system after a server failure...
Correct answer:
- the latest snapshot of the main store has to be loaded into main memory and logs have to be replayed
Incorrect answers:
- the latest snapshot has to be loaded and all users have to redo/restart the transactions that were lost
- the system can rely on the consistent state of data in main memory and only has to replay the latest log files
- only the caches have to be refilled, everything else is still available in main memory
General explanation: To recover an in-memory database, the latest snapshot of the main store has to be loaded into main memory and the additional logs have to be replayed. Since main memory is volatile, there is no consistent state in main memory and the system can not rely on that. Also, users do not have to redo their transactions.
Question Name: Transactions in Replicas
Text: What is important for transactions in a replica?
Correct answer:
- If a transaction with the ID number 'n' is visible, all transactions that committed before that transaction need to be also visible
Incorrect answers:
- If a transction with the ID number 'n' is visible, all transaction that have higher IDs have to be also visible
- It is important that the transaction IDs differ for each replica in order to ensure a coherent dataset when stitching the parts back together
- Each transaction might be replicated only once
General explanation: The highest transaction number of a replica represents the timeliness of its data and guarantees the completeness. If a replica receives a transaction that has an ID higher than its currently highest ID + 1, it is certain that the replica missed some messages. Before accepting the incoming transaction, the replica will therefore request the missing transactions from the master (or other replicas) before commencing. This mechanism would not be possible with different or unique IDs, thereby rendering the other answers as wrong.
Question Name: k-safety
Text: What does k-safety mean?
Correct answer:
- k-safety means that the master waits for k individual replica nodes acknowledge to have received the changes before acknowledging a successful write to the client
Incorrect answers:
- k-safety is the successor of j-safety, a standard defining access control levels in enterprise systems
- k-safety means that we replicate data to k systems in total
- k-safety means that we have to read all data changes at least k-times from the replicas to be certain that they equal each other
General explanation: k-safety means that the master waits for k individual replica nodes acknowledge to have received the changes before acknowledging a successful write to the client. The replicas do not have to acknowledge that they have persisted the changes, it is assumed that they will handle that or request the data again if something went wrong. By just acknowledging the receipt of the data, the response time gets considerably shorter. The other answers are made up and simply wrong.
Question Name: Scheduling
Text: Scheduling in a database context is...
Correct answer:
- a strategy to assign the computer's resources such as processing power, memory, or network bandwidth to running queries
Incorrect answers:
- a strategy to filter out nonrelevant results of database queries
- a strategy to produce more exact query results
- a strategy to format the results of database queries for presentation
General explanation: Scheduling in general is the allocation of available resources to processes. In a database context, these processes are usually queries or stored procedures. The other answers are wrong.
Question Name: Transactional Query Response Times
Text: The response time for transactional workloads...
Correct answer:
- has to be guaranteed, even for peak-load situations
Incorrect answers:
- is always so fast that it is negligible
- is of no importance
- is always much slower than that of analytical queries
General explanation: Response times are always important and unfortunately they are not always negligible for transactional workloads. Usually, transactional queries are much faster than analytical queries, but it is additionally required to guarantee some maximal response times in order to rule out negative effects to business processes. For a query that is needed to conclude the check out in a shopping process, it is essential to guarantee a response time even under peak loads to enable efficient marketing. If the system can not handle peak loads that are induced by promotions, profit is would be lost.
Question Name: Layers and Tiers
Text: The differentiation between tier and layer is that ...
Correct answer:
- tier is used for physical systems, whereas layer is used to descibe logical separations.
Incorrect answers:
- tier is used for logical separations, whereas layer is used to describe physical systems.
- not existent.
- tier was mainly used during the 80s and has been replaced by the word layer in the late 90s.
General explanation: This is simply a clarification of the word usage. Despite the difference is not always pointed out in all texts, it is clearly existent. Layer is the word to describe a logical separation, the word tier is used to describe different physical machines. Since these concepts are timeless and do not rely on any hardware or software evolution, the word usage also has not changed during the latest 30 years.
Question Name: Real Customer Data
Text: Real customer data should be used during development, because ...
Correct answer:
- it reflects actual workloads and use cases which helps to pinpoint realistic needs
Incorrect answers:
- the corresponding generated test data is orders of magnitudes bigger than the real data and would not fit into memory
- it can be sold to competitors if the customer does not pay for the developed software
- it is much less complex than generated data
General explanation: Real customer data should always be used if available. It reflects actual workloads and often contains additional valueable information like incomplete or inconsistent datasets. This leads to an improved stability of the software in development. If no real data is available, test data with realisitic characteristics concerning the amount of entries, distinct values and the existing distribution should be used. This test data is therefore not significantly bigger than the actual data, but It has the downside that it is often less complex, which means it lacks certain errors or unexpected values. Of course, the answer that the data could be used as a pressurizing medium is also wrong, since this is clearly illegal.