• Peter Zijlstra's avatar
    sched/stop_machine: Fix deadlock between multiple stop_two_cpus() · b17718d0
    Peter Zijlstra authored
    Jiri reported a machine stuck in multi_cpu_stop() with
    migrate_swap_stop() as function and with the following src,dst cpu
    pairs: {11,  4} {13, 11} { 4, 13}
    
                            4       11      13
    
    cpuM: queue(4 ,13)
                            *Ma
    cpuN: queue(13,11)
                                    *N      Na
                            *M              Mb
    cpuO: queue(11, 4)
                            *O      Oa
                                    *Nb
                            *Ob
    
    Where *X denotes the cpu running the queueing of cpu-X and X[ab] denotes
    the first/second queued work.
    
    You'll observe the top of the workqueue for each cpu: 4,11,13 to be work
    from cpus: M, O, N resp. IOW. deadlock.
    
    Do away with the queueing trickery and introduce lg_double_lock() to
    lock both CPUs and fully serialize the stop_two_cpus() callers instead
    of the partial (and buggy) serialization we have now.
    Reported-by: default avatarJiri Olsa <jolsa@redhat.com>
    Signed-off-by: default avatarPeter Zijlstra (Intel) <peterz@infradead.org>
    Cc: Andrew Morton <akpm@linux-foundation.org>
    Cc: Borislav Petkov <bp@alien8.de>
    Cc: H. Peter Anvin <hpa@zytor.com>
    Cc: Linus Torvalds <torvalds@linux-foundation.org>
    Cc: Oleg Nesterov <oleg@redhat.com>
    Cc: Peter Zijlstra <peterz@infradead.org>
    Cc: Rik van Riel <riel@redhat.com>
    Cc: Thomas Gleixner <tglx@linutronix.de>
    Link: http://lkml.kernel.org/r/20150605153023.GH19282@twins.programming.kicks-ass.netSigned-off-by: default avatarIngo Molnar <mingo@kernel.org>
    b17718d0
lglock.c 2.5 KB