FirstHack Learn
Log in Sign up free
Lessons in this course 0/6 All courses Operating Systems

CSE

Progress0 / 6 lessons
  1. 1. What an OS actually does, and what a system call is
  2. 2. Processes vs threads, with a fork() example
  3. 3. CPU scheduling with worked FCFS, SJF and Round Robin numbers
  4. 4. Deadlock: the four conditions and how to break them
  5. 5. Memory: paging, virtual memory and page faults
  6. 6. File systems and how a file is really stored

Courses › Operating Systems

Processes vs threads, with a fork() example

What a process owns, what a thread shares, and why fork() returns twice.

12 min read · Lesson 2 of 6 · Free

A process is a program plus its private world

A program is a file on disk. A process is that program running, together with everything it owns:

  • its own address space — code, data, heap, stack
  • its open file descriptors
  • its process ID, its parent, its user
  • its current CPU register values when it is not running

Two processes running the same program have completely separate memory. Neither can see the other's variables. That is the isolation the OS promised.

fork() makes a copy

fork() creates a new process that is a near-exact copy of the caller. Same code, same variable values, same open files. The two are then independent.

The famous strange part: fork() returns twice. Once in the parent, once in the child. That is not magic — there are now two processes, and each one is sitting at the same instruction, each about to receive a return value. The values differ so the two copies can tell themselves apart.

C
#include <stdio.h>
#include <unistd.h>
#include <sys/wait.h>

int main(void) {
    int x = 10;
    pid_t pid = fork();

    if (pid < 0) {
        perror("fork failed");
        return 1;
    }

    if (pid == 0) {
        x = x + 5;
        printf("child : pid=%d  x=%d\n", (int)getpid(), x);
    } else {
        wait(NULL);
        x = x + 100;
        printf("parent: pid=%d  x=%d  child was %d\n",
               (int)getpid(), x, (int)pid);
    }
    return 0;
}

Compile with gcc fork_demo.c -o fork_demo and run it. Output looks like:

Code
child : pid=4813  x=15
parent: pid=4812  x=110  child was 4813

Three things to take from that output.

In the child, fork() returned 0. In the parent it returned the child's PID. That is the whole protocol — 0 means "you are the child", a positive number means "you are the parent, and here is your child's ID", negative means failure.

x is 15 in one and 110 in the other. Both started from 10. The child's x = x + 5 had no effect on the parent's copy, and the parent's x = x + 100 had no effect on the child's. Separate address spaces, separate variables. This is the single most tested fact about fork().

wait(NULL) makes the parent pause until the child finishes. Without it, the parent might print first, or exit while the child is still running, leaving the child parented to init.

⚠️

A very common exam trap: how many processes does fork(); fork(); fork(); create in total? Each fork doubles the process count, so you end with 2 x 2 x 2 = 8 processes, of which 7 are new children. Draw the tree; do not guess.

The copy is not physically made straight away. Linux uses copy-on-write: parent and child share the same physical pages, marked read-only, and a page is duplicated only when one of them writes to it. That is why fork() on a 2 GB process is fast.

Threads share the address space

A thread is a separate line of execution inside one process. Threads in the same process share:

  • the whole address space — globals, heap, everything
  • open file descriptors

Each thread has only its own stack, its own registers, and its own program counter.

C
#include <stdio.h>
#include <pthread.h>

int counter = 0;

void *bump(void *arg) {
    (void)arg;
    for (int i = 0; i < 100000; i++) {
        counter++;
    }
    return NULL;
}

int main(void) {
    pthread_t t1, t2;
    pthread_create(&t1, NULL, bump, NULL);
    pthread_create(&t2, NULL, bump, NULL);
    pthread_join(t1, NULL);
    pthread_join(t2, NULL);
    printf("counter = %d (expected 200000)\n", counter);
    return 0;
}

Compile with gcc race.c -o race -pthread and run it several times. You will often get something like 137412 or 168903, not 200000, and a different number each run.

Why? counter++ is not one machine instruction. It is three: load counter into a register, add 1, store it back. Two threads can both load 5, both compute 6, and both store 6. Two increments, one net change. That is a race condition.

The fix is a mutex:

C
#include <stdio.h>
#include <pthread.h>

int counter = 0;
pthread_mutex_t lock = PTHREAD_MUTEX_INITIALIZER;

void *bump(void *arg) {
    (void)arg;
    for (int i = 0; i < 100000; i++) {
        pthread_mutex_lock(&lock);
        counter++;
        pthread_mutex_unlock(&lock);
    }
    return NULL;
}

int main(void) {
    pthread_t t1, t2;
    pthread_create(&t1, NULL, bump, NULL);
    pthread_create(&t2, NULL, bump, NULL);
    pthread_join(t1, NULL);
    pthread_join(t2, NULL);
    printf("counter = %d\n", counter);
    return 0;
}

Now it prints 200000 every single time.

Choosing between them

Process Thread
Memory private shared
Creation cost higher lower
Crash of one others survive whole process dies
Communication needs pipes, sockets, shared memory just use a variable
Bug risk isolation protects you races and deadlocks

Threads are faster to create and easy to communicate between, and that ease is exactly what makes them dangerous. Chrome uses a separate process per tab on purpose — a crashing page should not kill the browser.

💡

For the viva, be ready with one sentence: "A process has its own address space; threads of one process share it. That single difference explains creation cost, communication style and crash behaviour."