Commits · 73e75b416ffcfa3a84952d8e389a0eca080f00e1 · Kirill Smelkov / linux

31 Dec, 2008 40 commits

KVM: ppc: Implement in-kernel exit timing statistics · 73e75b41

Hollis Blanchard authored Dec 02, 2008

Existing KVM statistics are either just counters (kvm_stat) reported for
KVM generally or trace based aproaches like kvm_trace.
For KVM on powerpc we had the need to track the timings of the different exit
types. While this could be achieved parsing data created with a kvm_trace
extension this adds too much overhead (at least on embedded PowerPC) slowing
down the workloads we wanted to measure.

Therefore this patch adds a in-kernel exit timing statistic to the powerpc kvm
code. These statistic is available per vm&vcpu under the kvm debugfs directory.
As this statistic is low, but still some overhead it can be enabled via a
.config entry and should be off by default.

Since this patch touched all powerpc kvm_stat code anyway this code is now
merged and simplified together with the exit timing statistic code (still
working with exit timing disabled in .config).
Signed-off-by: Christian Ehrhardt <ehrhardt@linux.vnet.ibm.com>
Signed-off-by: Hollis Blanchard <hollisb@us.ibm.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

73e75b41

KVM: ppc: save and restore guest mappings on context switch · c5fbdffb

Hollis Blanchard authored Dec 02, 2008

Store shadow TLB entries in memory, but only use it on host context switch
(instead of every guest entry). This improves performance for most workloads on
440 by reducing the guest TLB miss rate.
Signed-off-by: Hollis Blanchard <hollisb@us.ibm.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

c5fbdffb

KVM: ppc: directly insert shadow mappings into the hardware TLB · 7924bd41

Hollis Blanchard authored Dec 02, 2008

Formerly, we used to maintain a per-vcpu shadow TLB and on every entry to the
guest would load this array into the hardware TLB. This consumed 1280 bytes of
memory (64 entries of 16 bytes plus a struct page pointer each), and also
required some assembly to loop over the array on every entry.

Instead of saving a copy in memory, we can just store shadow mappings directly
into the hardware TLB, accepting that the host kernel will clobber these as
part of the normal 440 TLB round robin. When we do that we need less than half
the memory, and we have decreased the exit handling time for all guest exits,
at the cost of increased number of TLB misses because the host overwrites some
guest entries.

These savings will be increased on processors with larger TLBs or which
implement intelligent flush instructions like tlbivax (which will avoid the
need to walk arrays in software).

In addition to that and to the code simplification, we have a greater chance of
leaving other host userspace mappings in the TLB, instead of forcing all
subsequent tasks to re-fault all their mappings.
Signed-off-by: Hollis Blanchard <hollisb@us.ibm.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

7924bd41

powerpc/44x: declare tlb_44x_index for use in C code · c0ca609c

Hollis Blanchard authored Dec 02, 2008

KVM currently ignores the host's round robin TLB eviction selection, instead
maintaining its own TLB state and its own round robin index. However, by
participating in the normal 44x TLB selection, we can drop the alternate TLB
processing in KVM. This results in a significant performance improvement,
since that processing currently must be done on *every* guest exit.

Accordingly, KVM needs to be able to access and increment tlb_44x_index.
(KVM on 440 cannot be a module, so there is no need to export this symbol.)
Signed-off-by: Hollis Blanchard <hollisb@us.ibm.com>
Acked-by: Josh Boyer <jwboyer@linux.vnet.ibm.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

c0ca609c

KVM: ppc: support large host pages · 89168618

Hollis Blanchard authored Dec 02, 2008

KVM on 440 has always been able to handle large guest mappings with 4K host
pages -- we must, since the guest kernel uses 256MB mappings.

This patch makes KVM work when the host has large pages too (tested with 64K).
Signed-off-by: Hollis Blanchard <hollisb@us.ibm.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

89168618

KVM: split out kvm_free_assigned_irq() · 4a643be8

Mark McLoughlin authored Dec 01, 2008

Split out the logic corresponding to undoing assign_irq() and
clean it up a bit.
Signed-off-by: Mark McLoughlin <markmc@redhat.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

4a643be8

KVM: add KVM_USERSPACE_IRQ_SOURCE_ID assertions · 61552367

Mark McLoughlin authored Dec 01, 2008

Make sure kvm_request_irq_source_id() never returns
KVM_USERSPACE_IRQ_SOURCE_ID.

Likewise, check that kvm_free_irq_source_id() never accepts
KVM_USERSPACE_IRQ_SOURCE_ID.
Signed-off-by: Mark McLoughlin <markmc@redhat.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

61552367

KVM: don't free an unallocated irq source id · f29b2673

Mark McLoughlin authored Dec 01, 2008

Set assigned_dev->irq_source_id to -1 so that we can avoid freeing
a source ID which we never allocated.
Signed-off-by: Mark McLoughlin <markmc@redhat.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

f29b2673

KVM: make kvm_unregister_irq_ack_notifier() safe · fdd897e6

Mark McLoughlin authored Dec 01, 2008

We never pass a NULL notifier pointer here, but we may well
pass a notifier struct which hasn't previously been
registered.

Guard against this by using hlist_del_init() which will
not do anything if the node hasn't been added to the list
and, when removing the node, will ensure that a subsequent
call to hlist_del_init() will be fine too.

Fixes an oops seen when an assigned device is freed before
and IRQ is assigned to it.
Signed-off-by: Mark McLoughlin <markmc@redhat.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

fdd897e6

KVM: remove the IRQ ACK notifier assertions · 844c7a9f

Mark McLoughlin authored Dec 01, 2008

We will obviously never pass a NULL struct kvm_irq_ack_notifier* to
this functions. They are always embedded in the assigned device
structure, so the assertion add nothing.

The irqchip_in_kernel() assertion is very out of place - clearly
this little abstraction needs to know nothing about the upper
layer details.
Signed-off-by: Mark McLoughlin <markmc@redhat.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

844c7a9f

KVM: VMX: fix sparse warning · efff9e53

Hannes Eder authored Nov 28, 2008

Impact: make global function static

  arch/x86/kvm/vmx.c:134:3: warning: symbol 'vmx_capability' was not declared. Should it be static?
Signed-off-by: Hannes Eder <hannes@hanneseder.net>
Signed-off-by: Avi Kivity <avi@redhat.com>

efff9e53

KVM: fix sparse warning · e8ba5d31

Hannes Eder authored Nov 28, 2008

Impact: make global function static

  virt/kvm/kvm_main.c:85:6: warning: symbol 'kvm_rebooting' was not declared. Should it be static?
Signed-off-by: Hannes Eder <hannes@hanneseder.net>
Signed-off-by: Avi Kivity <avi@redhat.com>

e8ba5d31

KVM: Remove extraneous semicolon after do/while · f3fd92fb
Avi Kivity authored Nov 29, 2008
```
Notices by Guillaume Thouvenin.
Signed-off-by: Avi Kivity <avi@redhat.com>
```
f3fd92fb

KVM: x86 emulator: fix popf emulation · 2b48cc75

Avi Kivity authored Nov 29, 2008

Set operand type and size to get correct writeback behavior.
Signed-off-by: Avi Kivity <avi@redhat.com>

2b48cc75

KVM: x86 emulator: fix ret emulation · cf5de4f8

Avi Kivity authored Nov 28, 2008

'ret' did not set the operand type or size for the destination, so
writeback ignored it.
Signed-off-by: Avi Kivity <avi@redhat.com>

cf5de4f8

KVM: x86 emulator: switch 'pop reg' instruction to emulate_pop() · 8a09b687
Avi Kivity authored Nov 27, 2008
```
Signed-off-by: Avi Kivity <avi@redhat.com>
```
8a09b687
KVM: x86 emulator: allow pop from mmio · 781d0edc
Avi Kivity authored Nov 27, 2008
```
Signed-off-by: Avi Kivity <avi@redhat.com>
```
781d0edc
KVM: x86 emulator: Extract 'pop' sequence into a function · faa5a3ae
Avi Kivity authored Nov 27, 2008
```
Switch 'pop r/m' instruction to use the new function.
Signed-off-by: Avi Kivity <avi@redhat.com>
```
faa5a3ae

KVM: Prevent trace call into unloaded module text · b8209182

Wu Fengguang authored Nov 26, 2008

Add marker_synchronize_unregister() before module unloading.
This prevents possible trace calls into unloaded module text.
Signed-off-by: Wu Fengguang <wfg@linux.intel.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

b8209182

KVM: s390: Fix memory leak of vcpu->run · 6692cef3

Christian Borntraeger authored Nov 26, 2008

The s390 backend of kvm never calls kvm_vcpu_uninit. This causes
a memory leak of vcpu->run pages.
Lets call kvm_vcpu_uninit in kvm_arch_vcpu_destroy to free
the vcpu->run.
Signed-off-by: Christian Borntraeger <borntraeger@de.ibm.com>
Acked-by: Carsten Otte <cotte@de.ibm.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

6692cef3

KVM: s390: Fix refcounting and allow module unload · d329c035

Christian Borntraeger authored Nov 26, 2008

Currently it is impossible to unload the kvm module on s390.
This patch fixes kvm_arch_destroy_vm to release all cpus.
This make it possible to unload the module.

In addition we stop messing with the module refcount in arch code.
Signed-off-by: Christian Borntraeger <borntraeger@de.ibm.com>
Acked-by: Carsten Otte <cotte@de.ibm.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

d329c035

KVM: x86 emulator: consolidate emulation of two operand instructions · 6b7ad61f
Avi Kivity authored Nov 26, 2008
```
No need to repeat the same assembly block over and over.
Signed-off-by: Avi Kivity <avi@redhat.com>
```
6b7ad61f
KVM: x86 emulator: reduce duplication in one operand emulation thunks · dda96d8f
Avi Kivity authored Nov 26, 2008
```
Signed-off-by: Avi Kivity <avi@redhat.com>
```
dda96d8f

KVM: MMU: optimize set_spte for page sync · ecc5589f

Marcelo Tosatti authored Nov 25, 2008

The write protect verification in set_spte is unnecessary for page sync.

Its guaranteed that, if the unsync spte was writable, the target page
does not have a write protected shadow (if it had, the spte would have
been write protected under mmu_lock by rmap_write_protect before).

Same reasoning applies to mark_page_dirty: the gfn has been marked as
dirty via the pagefault path.

The cost of hash table and memslot lookups are quite significant if the
workload is pagetable write intensive resulting in increased mmu_lock
contention.
Signed-off-by: Marcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: Avi Kivity <avi@redhat.com>