Python Process Lifecycle: Using multiprocessing start() and join()
Learn how to start and join Python multiprocessing processes, use timeouts, clean up workers, and avoid common lifecycle pitfalls such as zombies and deadlocks.
The multiprocessing module's Process class lets you manage the lifecycle of a child process. start() begins execution in a separate OS process, and join() blocks the caller until that process terminates. Knowing what each method does, and what it does not do, is important for writing reliable concurrent Python programs.
The Role of start() and join() in Python Processes
multiprocessing.Process represents a separate OS process. Unlike threads, processes have their own memory space, so they are not constrained by Python's GIL and can run in parallel on multi-core systems. start() launches the process and returns immediately; the target function runs asynchronously. join() makes the calling process wait until the target process has finished. Without join(), the parent can exit before the child completes, leaving a child process running without coordination from the rest of the program.
Creating and Starting a Process with multiprocessing
Here is the minimal pattern:
import multiprocessing import time def worker(name): print(f'Worker {name} starting') time.sleep(2) print(f'Worker {name} finished') if __name__ == '__main__': p = multiprocessing.Process(target=worker, args=('A',)) p.start() print('Process started') p.join() print('Process joined')
The if __name__ == '__main__' guard is mandatory on Windows and recommended on all platforms to avoid recursive process creation when the module is imported. start() returns without waiting for the worker to finish, so 'Process started' may appear before the worker prints. join() blocks until the worker finishes, so 'Process joined' appears after the worker's final message.
What join() Actually Waits For
By default, join() waits indefinitely for the process to terminate. It does not return the process's exit code; check p.exitcode after the process has finished. join() also accepts an optional timeout argument, measured in seconds. If the process does not terminate within the timeout, join() returns anyway, and the process continues running in the background. You can then check p.is_alive() to see whether it is still running.
p.join(timeout=1.0) if p.is_alive(): print('Process still running, will terminate it') p.terminate() p.join()
terminate() sends a termination request to the child. After termination, call join() again to wait for cleanup and release the process resource. Without that second join(), a finished child can remain unreaped, and on POSIX systems may become a zombie until the parent exits.
Handling Process Timeouts and Cleanup
A common pattern is to give a process a fixed amount of time to complete, then terminate it if it exceeds that limit. This is useful for tasks that may hang due to I/O or external dependencies. join(timeout) is the primary tool for this. After the timeout, decide whether to let the process continue or terminate it. If you terminate it, call join() again to reap the process.
p = multiprocessing.Process(target=long_running_task) p.start() p.join(5) if p.is_alive(): p.terminate() p.join() print('Task terminated after timeout') else: print('Task completed within 5 seconds')
Common Pitfalls: Deadlocks, Zombies, and Shared State
One subtle issue with join() is deadlock. If a child process waits for input from the parent, and the parent is blocked in join() waiting for the child, neither can proceed. This often happens when a child writes to a pipe or queue that the parent never reads: the child blocks on a full buffer, and the parent blocks in join(). To avoid this, read from pipes and queues that children write to, or drain them promptly.
Another pitfall is forgetting to call join(). If a child finishes and the parent never reaps it, the child can become a zombie on POSIX systems until the parent joins it or exits. If the parent exits while the child is still running, the child can become an orphan and keep running independently.
Shared state between processes is another source of confusion. Because processes have separate memory spaces, a global variable in the parent is not automatically visible in the child. Passing mutable objects through Process arguments works only if they are picklable and are copied, not shared. For true sharing, use multiprocessing.Value, Array, or a Manager.
Daemon Processes and Their Interaction with join()
A process can be marked as a daemon by setting daemon=True before calling start(). Daemon processes are terminated automatically when the main process exits, so they are useful for background workers that should not outlive the parent. Calling join() on a daemon process is allowed, but daemon=True does not guarantee the process will finish before shutdown; if the daemon is still running when the main process exits, it is terminated automatically. If you need a background task to finish before the program exits, use a non-daemon process and call join() explicitly.
p = multiprocessing.Process(target=worker, daemon=True) p.start() # Main process may exit, killing the daemon
When to Use Process Pools Instead of Manual start/join
Managing individual processes with start() and join() is appropriate when you need fine-grained control over each process's lifecycle, such as different timeouts or termination policies. For many workloads, however, multiprocessing.Pool is simpler and more robust. A pool manages a fixed number of worker processes and distributes tasks automatically. You submit tasks with apply_async or map, and the pool manages the workers for you. When managing a pool manually, call close() to stop accepting tasks and join() to wait for all workers to finish.
from multiprocessing import Pool def square(x): return x * x with Pool(4) as pool: results = pool.map(square, range(10)) print(results)
Pool.map() blocks until all results are ready, so by the time the with block exits the work has already finished. The pool's context manager then shuts the workers down. Use a pool when the number of tasks is large and the tasks are independent. Use manual Process objects when you need per-process control: for example, a long-running service that must be terminated on demand, or a small number of processes with different arguments and timeouts.
When you use manual processes, remember that join() is part of cleanup. Even if you do not need the result, pairing every start() with a join()—directly or after terminate()—ensures that the process's resources are released. In long-running parent processes, failing to join completed children can accumulate zombie processes and exhaust system resources.