Keir Fraser [Mon, 2 Mar 2009 10:26:37 +0000 (10:26 +0000)]
hvm: passthrough MSI-X mask bit acceleration
Add a new parameter to DOMCTL_bind_pt_irq to allow Xen to know the
guest physical address of MSI-X table. Also add a new MMIO intercept
handler to intercept that gpa in order to handle MSI-X vector mask
bit operation in the hypervisor. This reduces the load of device model
considerably if the guest does mask and unmask frequently
Keir Fraser [Mon, 2 Mar 2009 10:23:50 +0000 (10:23 +0000)]
xend: Fix some mistakes in tools/python/xend/util/pci.py
1) PCI_PM_CTRL_NO_SOFT_RESET: this is bit3 of PMCSR(Power Management
Control/Status). It should be 8. This bit means a device's capability
of not doing an internal reset across D3hot/D0.
If the bit is 1, there shall be no reset across D3hot/D0, so we should
not use it as a method to reset device.
2) When performing reset using standard FLR methods, we should sleep
at least 100ms; in current code, it's incorrect somewhere we sleep
0.2s and somewhere 0.01s.
3) In detect_dev_info(), fix a typo: PCI_EXP_TYPE_PCI_BRIDG ->
PCI_EXP_FLAGS_TYPE.
4) fix a small typo in the comment of transform_list().
Keir Fraser [Sun, 1 Mar 2009 14:58:07 +0000 (14:58 +0000)]
x86, hvm: gcc44 build fix.
Broken constrain in inline asm. Bytewise access works with a, b, c, d
registers only, thus "r" is wrong, it must be "q". gcc 4.4 tries to
use the si register, which doesn't work and thus fails the build.
Keir Fraser [Sun, 1 Mar 2009 14:50:04 +0000 (14:50 +0000)]
xenstored: fix use-after free bug
Problem: Handling requests for one connection can not only zap the
connection itself, due to socket disconnects for example. It can also
zap *other* connections, due to domain release requests. Especially
it can zap the connection we have saved a pointer to in the "next"
variable.
Keir Fraser [Fri, 20 Feb 2009 17:02:36 +0000 (17:02 +0000)]
xenconsole: Fix pty handling
I printed the terminal attributes after openpty() and they were
garbage on the first console, valid on the second etc.
openpty() gets garbage in (uninitialized attributes MODIFIED by
cfmakeraw()). It sets the slave to the attributes requested. Using
uninitialized data for cfmakeraw->openpty results in pty attributes
that may even have the receiver disabled. Closing the slave just hides
the bug as these attributes disappear and hope the slave will be
reopened and initialized.
From: Juergen Hannken-Illjes <hannken@netbsd.org> Signed-off-by: Christoph Egger <Christoph.Egger@amd.com>
Keir Fraser [Fri, 20 Feb 2009 11:11:40 +0000 (11:11 +0000)]
[VTD] Utilise the snoop control capability in shadow with VT-d code
We compute the shadow PAT index in leaf page entries now as:
1) No VT-d assigned: let shadow PAT index as WB, handled already
in shadow code before.
2) direct assigned MMIO area: let shadow code compute the shadow
PAT with gMTRR=UC and gPAT value.
3) Snoop control enable: let shadow PAT index as WB.
4) Snoop control disable: let shadow code compute the shadow
PAT with gMTRR and gPAT, handled already in shadow code before
Keir Fraser [Thu, 19 Feb 2009 10:59:43 +0000 (10:59 +0000)]
vt-d: workaround for Mobile Series 4 Chipset
Incorporated VT-d workaround for a sighting on Intel Mobile Series 4
chipset found in Linux iommu. The sighting is the chipset is not
reporting write buffer flush capability correctly.
Keir Fraser [Tue, 17 Feb 2009 11:10:00 +0000 (11:10 +0000)]
ia64: Enhance vt-d support for ia64.
This patch targets for enhancing vt-d support for ia64.
1. reserve enough memory for building dom0 vt-d page table.
2. build 1:1 vt-d page table according to system's mem map.
3. enable vt-d interrupt support for ia64.
Keir Fraser [Tue, 17 Feb 2009 11:06:16 +0000 (11:06 +0000)]
passthrough: fix MSI-X table fixmap allocation
Currently, msix table pages are allocated a fixmap page per vector,
the available fixmap pages will be depleted when assigning devices
with large number of vectors. This patch fixes it, and a bug that
prevents cross-page MSI-X table from working properly
It now allocates msix table fixmap pages per device, if the table
entries of two msix vectors share the same page, it will only be
mapped to fixmap once. A ref count is maintained so that it can
be unmapped when all the vectors are freed.
Also changes the meaning of msi_desc->mask_base from the va of msix
table start to the va of the target entry. The former one is currently
buggy (it always maps the first page but msix can support up to 2048
entries) and can't handle separately allocated pages.
Keir Fraser [Fri, 13 Feb 2009 09:38:16 +0000 (09:38 +0000)]
xenapi: Correct some syntax errors in xen/xend/XendAPI.py
- usage of undefined variables in error cases (invalid handle
specified) in methods VBD_create, VTPM_destroy, event_unregister
- not imported module 'uuid' in method debug_create results in an
exception
Keir Fraser [Fri, 13 Feb 2009 09:32:02 +0000 (09:32 +0000)]
xendomains: clean up output formatting
Show errors in the way they are coming from xm command only (no usage
is printed now). Watchdog_wm() has been changed for not showing dots
in the process of shutting down domains and if an error occurs it
prints target domain, operation (save/restore/migrate etc.) and reason
of failure in more user-friendly way.
Signed-off-by: Michal Novotny <minovotn@redhat.com>
Isaku Yamahata [Fri, 13 Feb 2009 02:23:16 +0000 (11:23 +0900)]
[IA64] shrink ia64 struct page_info.
This patch is the ia64 counter part of 19107:0858f961c77a,
19132:5848b49b74fc and 19136:162cdb596b9a.
This patch shrink ia64 struct page_info and rearrange its members.
The shrinking is made compile time option in config.h with default off
becuase physical address size is architected to 50bit by ia64
and mfn isn't always addressed by 32bits with 16KB page size.
Isaku Yamahata [Fri, 13 Feb 2009 01:56:01 +0000 (10:56 +0900)]
[IA64] MCA: Avoid calling xmcalloc from interrupt handler
This patch fixes to avoid calling xmalloc() from the interrupt handler.
Calling xmalloc() with interrupt disabled triggers the following
BUG_ON().
> (XEN) Xen BUG at xmalloc_tlsf.c:548
Keir Fraser [Thu, 12 Feb 2009 10:48:55 +0000 (10:48 +0000)]
Cleanup naming for ia64 and x86 interrupt handling functions
- Append '_IRQ' to AUTO_ASSIGN, NEVER_ASSIGN, and FREE_TO_ASSIGN
- Rename {request,setup}_irq to {request,setup}_irq_vector
- Rename free_irq to release_irq_vector
- Add {request,setup,release}_irq wrappers for their
{request,setup,release}_irq_vector counterparts
- Added generic irq_to_vector inline for ia64
- Changed ia64 to use the new naming scheme
Keir Fraser [Wed, 11 Feb 2009 13:07:45 +0000 (13:07 +0000)]
x86: cpufreq get_cur_val adjustment
c/s 19149 update cpufreq get_cur_val logic to avoid cross processor
call. However, to avoid null drv_data pointer, we adjust some logic in
this patch to keep advantage of c/s 19149 and at same time to avoid
null drv_data pointer.
Keir Fraser [Wed, 11 Feb 2009 10:47:24 +0000 (10:47 +0000)]
cpufreq cmdline handling
c/s 19147 adjust cpufreq cmdline handling, this patch is a complement
to c/s 19147.
In this patch:
1. add common para (governor independent para) handling;
2. change governor dependent para handling method, governor dependent
para will only be handled by the handler of that governor (not by all
governors);
3. add userspace governor dependent para handling;
4. change para name 'threshold' of ondemand governor to 'up_threshold'
since ondemand has only 'up_threshold', and conservative governor (will be
implemented later) has both 'up_threshold' and 'down_threshold';
5. change some coding style (c/s 19147, drivers/cpufreq/cpufreq.c) to
keep coordination with original drivers/cpufreq/cpufreq.c coding style;
(originally this file is ported from linux, we partly use linux coding style)
Keir Fraser [Tue, 10 Feb 2009 05:51:00 +0000 (05:51 +0000)]
x86: mce: Provide extended physical CPU info.
Provide extended physial CPU info for the sake of dom0 MCE handling.
This information includes <cpu,core,thread> info for all logical CPUs,
cpuid information from all of them, and initial MSR values for a few
MSRs that are important to MCE handling.
Signed-off-by: Frank van der Linden <Frank.Vanderlinden@Sun.COM>
Keir Fraser [Fri, 6 Feb 2009 10:42:26 +0000 (10:42 +0000)]
Better separate IOAPIC management from interrupt vector management
Don't automatically update ioapic_irq array when allocating vectors.
Only do so when actually allocating IOAPIC irqs. Also move some
IOAPIC specific defines to io_apic.c.
Keir Fraser [Fri, 6 Feb 2009 10:36:23 +0000 (10:36 +0000)]
Cleanup IOMMU interrupt setup
- Check for errors when allocating interrupt vectors
- Clean up if interrupt allocation failed
- Make sure that the allocated vector is not reused
Keir Fraser [Thu, 5 Feb 2009 12:17:08 +0000 (12:17 +0000)]
libxc support for the new partial-HVM-save domctl.
This includes making the pagetable walker in xc_pagetab.c behave
correctly for 32-bit and 64-bit HVM guests.
Keir Fraser [Thu, 5 Feb 2009 12:16:53 +0000 (12:16 +0000)]
Remove uses of DECLARE_BITMAP in the public HVM save format headers
and replace them with open-coded equivalents. DECLARE_BITMAP is
not exported to user-space consumers of the Xen headers.
Keir Fraser [Thu, 5 Feb 2009 12:14:09 +0000 (12:14 +0000)]
x86: recover pat value for bsp after S3 resume.
host pat is set to cover all memory types by Xen, which is
necessary to support guest mtrr/pat, especially when device
is passthroughed with VT-d. However pat on bsp is not=20
recovered which could make assigned device defunct after S3
resume
Keir Fraser [Thu, 5 Feb 2009 12:09:10 +0000 (12:09 +0000)]
Add a page_info flag to indicate whether free pages need a TLB flush
on next use.
Apart from teh small performance gain of this, my primary motivation
is to avoid TLB flushes very early in boot, when the system is not yet
properly set up for cross-TLB shootdowns.
Keir Fraser [Wed, 4 Feb 2009 15:29:51 +0000 (15:29 +0000)]
Remove cpumask for page_info struct.
This makes TLB flushing on page allocation more conservative, but the
flush clock should still save us most of the time (page freeing and
alloc'ing tends to happen in batches, and not necesasrily close
together). We could add some optimisations to the flush filter if this
does turn out to be a significant overhead for some (useful)
workloads.
Keir Fraser [Wed, 4 Feb 2009 15:08:46 +0000 (15:08 +0000)]
x86: Clean up PV guest LDT handling.
1. Do not touch deferred_ops in invalidate_shadow_ldt(), as we may
not always be in a context where deferred_ops is valid.
2. Protected the shadow LDT with a lock, now that mmu updates are not
protected by the per-domain lock.
Keir Fraser [Wed, 4 Feb 2009 12:01:47 +0000 (12:01 +0000)]
Eliminate some special page list accessors
Since page_list_move_tail(), page_list_splice_init(), and
page_list_is_eol() are only used by relinquish_memory(), and that
function can easily be changed to use more generic accessors, just
eliminate them altogether.
Keir Fraser [Tue, 3 Feb 2009 18:13:55 +0000 (18:13 +0000)]
x86: misc adjustments to acpi-cpufreq
Avoid the call to check_freq() by default, since that function may
spin up to 1ms on certain systems without indicating any kind of severe
failure. This matches similar behavior in Linux.
Avoid doing a cross processor call in get_cur_val() if the current CPU
has its bit set in the mask passed in. Also use the local variable
'cpu' consistently, allowing to remove another local variable.
Keir Fraser [Tue, 3 Feb 2009 18:13:22 +0000 (18:13 +0000)]
cpufreq: attach __exit to the (unused) cpufreq governor exit handlers
... in order to make them disappear from the final image. Of course
they could as well be removed altogether, but I assumed that whoever
added them had a reason to do so.
Keir Fraser [Tue, 3 Feb 2009 18:12:51 +0000 (18:12 +0000)]
Consolidate cpufreq cmdline handling
... by moving as much of the option processing into cpufreq code as is
possible, by folding the cpufreq_governor option into the cpufreq one
(the governor name, if any, must be specified as the first thing after
the separator following "cpufreq=xen"), and by allowing each
governor to have an option processing routine.
Keir Fraser [Tue, 3 Feb 2009 18:11:03 +0000 (18:11 +0000)]
x86: Relocate Multiboot structures where we know they will be
accessible. GRUB2 seems to like to stick them really high sometimes
(just below 4GB).
The 32-bit C code framework that this sets up can also be used for
other stuff in future:
* early cmdline parsing
* relocating multiboot modules so they too are guaranteed accessible
Its interaction with normal Xen start-of-day, and with the 16-bit
assembly trampoline, needs a bit of thought.