CS110 Final Exam Studying notes

CS110 Final Exam Studying notes Lab 5

The implementation of acquireAsReader acquires the stateLock (via the lock_guard) before it does anything else, and it doesn’t release the stateLock until the method exits. Why can’t the implementation be this instead?

the state lock is released too early. With state lock released, someone else could possibly acquire the  the writeLock as we are acquiring the readLock. This should not happen.

The implementation of acquireAsWriter acquires the stateLock before it does anything else and it releases the stateLock just before it acquires the readLock. Why can’t acquireAsWriter adopt the same approach as acquireAsReader and just hold onto stateLock until the method returns?

This would present the possibility of deadlock. We would be waiting to acquire the readLock with the stateLock locked; releasing the readLock requires the stateLock too.

Notice that we have a single release method instead of releaseAsReader and releaseAsWriter methods. How does the implementation know if the thread acquired the rwlock as a writer instead of a reader (assuming proper use of the class)?

it doesn’t need to care. We assume a thread will not call release() unless it has one of the locks. In this case, if writeState == Writing, only one thread may have any lock, so that must be me. On the other hand, if this thread has a lock and nobody has the writeLock, this thread must have a readLock. 

The implementation of release relies on notify_all in one place and notify_one in another. Why are those the correct versions of notify to call in each case?

In the notify_all case, there might be several threads waiting to read. Now that we are done writing, they are all allowed to read. However, if we just finished reading, at most one thread will now get to write. We don’t want them all to go after the write lock now. 

A thread that owns the lock as a reader might want to upgrade its ownership of the lock to that of a writer without releasing the lock first. Besides the fact that it’s a waste of time, what’s the advantage of not releasing the read lock before re-acquiring it as a writer, and how could be the implementation of acquireAsWriter be updated so it can be called after acquireAsReader without an intervening release call?

What if someone else cuts in? We could smoothly handle this case by passing some flag indicating this mode. In that case, we should decrement numReaders once we have the readLock but before hanging on readCond.

In Part 2, I used a mutex (briefly acquired by each waiting minstrel). This worked fine, but a  condition_variable_any would have been cleaner. 

Lab 6

  • What does it mean when we say that a process has a private address space

  • Each process may ask for memory within the universe of addresses, but it’s all getting mapped to physical memory reserved for that process.

  • What are the advantages of a private address space?

  • Security. You simply can’t ask for what’s not yours. 

  • What are the disadvantages?

  • You can’t share objects in memory. 

  • What programming directives have we used in prior assignments and discussion section handouts to circumvent address space privacy?

  • ERROR: I forgot about ptrace. We can peek inside child processes. mmap also 

  • ref() I think makes something persistent across addresses. 

  • Also, we can use things like file descriptors, which are references to the same thing which are copied from one process to another when forked. 

  • In what cases do the processes whose private address spaces are being publicized have any say in the matter?

  • ERROR: You can’t do much. System configuration can disable ptrace for some or all processes. But virtualization means you don’t really know what you’re getting with, for example, stdin and stdout. They might be sneakily-piped. 

  • When architecting a larger program like **farm** or **stsh** that relies on multiprocessing, what did we need to do to exchange information across process boundaries?

  • We used signals. We can also pipe file descriptors, which allow us to transfer information.

  • Can a process be used to execute multiple executables? Restated, can it **execvp** twice to run multiple programs?

  • This might be a semantic issue… Yes, a process can execvp, transform itself into a new process, and launch an executable which contains another invocation of execvp. But then is it still the same process?

  • Threads are often called lightweight processes. In what sense are they processes? And why the lightweight distinction?

  • They are lightweight because they all exist within the same process, and don’t need to set up their own memory spaces. They are processes in that a program using threads can (appear to) do several different things at once. 

  • Threads are often called virtual processes as well. In what sense are threads an example of virtualization?

  • Threads are virtualization in that they seem to implement affordances of processes without doing so.

  • ERROR: Virtualization specifically refers to many-to-one or one-to-many abstractions

  • Threads running within the same process all share the same address space. What are the advantages and disadvantages of allowing threads to access pretty much all of virtual memory?

  • Pro: threads can much more easily access global variables. Con: It’s easier to create race conditions. 

  • What are the advantages of leveraging the pid abstraction for thread ids?

  • You can reuse infrastructure. For example, do threads get in line at the scheduler just like processes?

  • What happens if you pass a thread id that isn’t a process id to **waitpid**?

  • Death.

  • ERROR: No, you can do this.

  • What happens if you pass a thread id to **sched_setaffinity**?

  • Death.

  • ERROR: No, you can do this. 

  • What are the advantages of requiring that a thread always be assigned to the same CPU?

  • You don’t have to maintain memory state across CPUs. 

  • Why might you prefer multithreading over multiprocessing if both are reasonably good options?

  • It’s cheaper and faster. You don’t need to set up address space for each one. Probably, each thread can also read from the same port?

  • Why might you prefer multiprocessing over multithreading if both are reasonably good options?

  • On a webserver, for example, a process could get terribly gummed up (DDOS). 

  • What happens if a thread within a larger process calls **fork**?

  • All the threads get copied over!

  • What happens if a thread within a larger process calls **execvp**?

  • Goodbye everybody. We’re doing something new now. 

  • Assume you know nothing about multiprocessing but needed to emulate the functionality of **farm**. How could you use multithreading to design an equally parallel solution without ever relying on multiprocessing?

  • Spin off a manager thread, and have that spin up new threads for each piece of work to be done. 

  • Inversely, how could you have implemented **aggregate** to rely on multiprocessing (without threading) to arrive at an equally parallel solution?

  • I suppose you could create a ProcessPool and swap it in for the ThreadPool.

  • Could multithreading have contributed to the implementation of stsh in any meaningful way? Why or why not?

  • Nope. You can’t execvp new executables in a thread without taking the whole process down. And you can’t run an executable without execvp’ing it.

  • Is **i--** thread safe? Why or why not?

  • No. it’s actually several commands: decrement and report. If some other i— happened concurrently, you might report a severely decremented value.

  • What’s the difference between a **mutex** and a **semaphore** with an initial value of 1? Can one be substituted for the other?

  • You could interchange them in some cases, if you’re careful.  However, the semaphore could still be signalled by folks who never waited first.

  • ERROR: Also, the thread that locks a mutex must be the one to unlock it. 

  • What is the **lock_guard** class used for, and why is it useful?

  • Oh, it’s great. Acquires a lock on a mutex, and releases the lock as the function returns. Sometimes it’s convenient so you don’t forget, and other times it’s essential because you don’t want to unlock until the return value goes out or the exception is thrown. 

  • What is busy waiting? Is it ever a good idea? Does your answer to the good-idea question depend on the number of CPUs your multithreaded application has access to?

  • Busy waiting is creating a while (true) loop to occupy the processor as a way to stall for time. I can’t think of a time when it’s a good idea, unless you’re trying to warm up the house. 

  • As it turns out, the semaphore’s constructor allows a negative number to be passed in, as with **semaphore** **s(-11)**. Identify a scenario where -11 might be a sensible initial value.

  • In this case, it can collect disparate processes. For example, maybe you’re downloading twelve files and don’t want to proceed until they all come home, each thread could signal when it finished.

  • What would the implementation of **semaphore::signal(size_t increase = 1)** need to look like if we wanted to allow a **semaphore**’s encapsulated value to be promoted by the **increase** amount? Note that **increase** defaults to 1, so that this version could just replace the standard **semaphore::signal** that’s officially exported by the **semaphore** abstraction.

  • We would need to change the check for whether to notify on the condition variable. 

  • What’s the multiprocessing equivalent of the **mutex**?

  • SIGTSTP

  • ERROR: Signal masks, because they isolate a process temporarily from interruption to prevent race conditions.  

  • What’s the multiprocessing equivalent of the **condition_variable_any**?

  • Perhaps a custom signal with an implementation on the other side with a WAITPID which checks a condition. 

  • ERROR: Sigsuspend. 

Lab 7

  • Briefly describe a simple ThreadPool test program that would have deadlocked had ThreadPool::worker called allDone.notify_one instead of allDone.notify_all.

  • If you have a bunch of threads all waiting on the same pool, only one would be notified. 

  • ERROR: Take this all the way; call join() on each thread.

  • Assume ThreadPool::worker gets through the call to allDone.notify_all and gets swapped off the processor immediately after the lock_guard is destroyed. Briefly describe a situation where a thread that called ThreadPool::wait still won’t advance past the allDone.wait call.

  • Maybe, before the waiting thread gets a turn, some other job gets scheduled on the threadpool. 

  • Had allDone.notify_all been called unconditionally (i.e. not just because outstandingThunkCount was zero), the ThreadPool would have still worked correctly. Why is this true, and why is the if test still the right thing to include?

  • Every notification will cause waiting condition_variable_any variables to acquire a lock and evaluate their thunks. We would still be ok because outstandingThunkCount would still be nonzero. But this is irritating and requires extra contention for a lock on allDone.

In the first server, no backlog would ever accumulate because requests are handled sequentially. Instead, the request would get no response at all until the server was ready to handle it. Depending on the timeout set for the request and other requests, it might not be very likely that the request would receive a response at all within the time limit. 

The second server will accept requests as they come in, with no wait time. But then handling will be delayed. By the 500th request, the backlog will be quite full, and incoming requests will be accepted and then dropped. Possibly the 500th request will get lucky and come in the moment another is pulled off the backlog; otherwise it will be accepted and then dropped. 

ERROR: I misinterpreted where the backlog sits. It’s the backlog of requests which await acceptance. Therefore, the first one will fill the backlog queue and the second one will be accepted and hang for a long time. 

  • Explain the differences between a pipe and a socket.

  • While pipes and sockets are both represented by file descriptors, a pipe connects processes on the same machine, while a socket represents a connection between processes potentially on different machines.  

  • ERROR: Pipe is unidirectional; socket is bidirectional. 

  • Explain how system calls are a form of client/server and request/response.

  • You could see system calls as a form of client/server in that syscalls populate registers with the “request” content and then are reactivated (signalled) as a response once the response value has been put on the register and desired kernel-side side effects have taken place. 

  • Describe how networking is just another form of function call and return. What “function” is being called? What are the parameters? And what’s the return value?

  • Networking is just function call and return in that we “call” a subprocess identified by a URL, port number, request headers, and request payload, pass arguments with the same, and then, either blocking or async, wait for the response, which is the return. HTTP error messages are similar to the -1 returned by subprocesses, not inherently raising exceptions, but still representing an unexpected state. 

  • Describe the network architecture needed for:

  • Email servers and clients

  • You have a network of servers handling mail, passing it around. In POP, your computer acts as the last link in the server chain, in IMAP you are more explicitly a client… DNS lookup, … 

  • Peer-to-Peer Text Messaging via cell phone numbers

  • You’ve got client phones and synchronized servers. 

  • Probably, there’s something like Apple’s notification service which bundles all app notifications together so that the client phones can just issue one ping to find out if they have notifications. 

  • Skype

  • You’ve got client programs on the computer, which open up sockets to the servers. 

  • Probably, none of this is truly peer-to-peer. 

  • As it turns out, each of the three network applications above all make use of custom protocols. Which ones could have relied on HTTP and/or HTTPS instead of custom protocols?

  • Email and text messaging would be ok. For skype, you have different priorities. If something doesn’t arrive in time, it missed its chance. Focus on sending what’s new. 

  • Note the very first line of the server’s main function creates a server socket that listens to 33334 for incoming network activity. Why, after the for loop within main, are all of the child processes also listening to port 33334 through the same server socket descriptor?

  • the socket descriptor is a file descriptor, which is copied on fork.

  • During lecture, we referred to port numbers as virtual process ids. What did we mean by that, and why are virtual process ids needed with networked applications?

  • Processes bind to ports; from the outside, we can communicate with particular processes on another machine by specifying the port. The ports are virtualized in that you don’t need to (and don’t want to) have to specify the arbitrary pid on the other machine. 

  • What happens if the call to exit(0) is removed from the implementation of shutdownServers and we send a SIGINT to the master process in the suite of 4 hello-server servers?

  • I don’t know.

  • ERROR: The parent process fails to exit. Why should it? 

  • If the call to the signal function is moved to reside above the for loop in main instead of below it, does that impact our ability to close down the full suite of 4 hello-server servers?

  • I think so. Signal handlers are not copied across a fork.

  • The above program doesn’t close the server sockets in shutdownServers. Describe a simple way to ensure a server socket is properly closed before the surrounding server executable exits.

  • I forget. 

  • ERROR: Oh. there was no trick. Just remember who you opened so you can close it at the end. in the SIGINT handler. 

**Stuff to put on cheat sheet: **

Stuff on cheat sheet EventBarrier solution code. (shows use of mutex, lock_guard, and condition_variable_any)

Maybe handleRequest in lab 7. Maybe solution to ThreadPool problem from exam 1

Solution to Problem 4, DNS Server. I’m not strong with 

Current grade summary: My grades are 83% on homework and 87% on midterm. I’d say the class average is probably a good bit above that.