Assorted 80386 tweaks
This commit is contained in:
parent
c9e303600d
commit
f5c0930f83
8 changed files with 255 additions and 232 deletions
|
|
@ -5,11 +5,9 @@ Assembling a detailed and accurate history of the 80386, including a complete li
|
|||
problems were fixed by a later stepping, seems virtually impossible at this late date.
|
||||
|
||||
I won't make the attempt here, either. Using information from various sources, I'll start with an overview
|
||||
of the steppings (revision levels), including how each stepping was externally marked and internally identified,
|
||||
along with lists of associated errata, then move on to more detailed errata information (from Intel's own documents),
|
||||
and end with a summary of undocumented 80386 instructions.
|
||||
|
||||
Until further notice, this document is a work-in-progress.
|
||||
of the steppings, including how each stepping was externally marked and internally identified, along with lists
|
||||
of associated errata, then move on to more detailed errata information (from Intel's own documents), and end
|
||||
with a summary of undocumented 80386 instructions.
|
||||
|
||||
### Steppings
|
||||
|
||||
|
|
@ -378,7 +376,6 @@ Here's what the world knew about 80386 problems in the B1 stepping, as of Decemb
|
|||
fields are correctly stored. Note that coprocessor error-handling routines are the only routines possibly
|
||||
affected. Also note that the problem does not occur in PROTECTED mode programs (since no opcode is saved
|
||||
by FSAVE or FSTENV in that case).
|
||||
|
||||
**Workaround**: In REAL mode or in VIRTUAL 8086 mode, the instruction linear address field can be used to
|
||||
read the opcode from memory. Note that the two bytes fetched need to be swapped to yield the image that
|
||||
FSAVE and FSTENV normally stores.
|
||||
|
|
@ -386,7 +383,6 @@ Here's what the world knew about 80386 problems in the B1 stepping, as of Decemb
|
|||
**Problem**: If either of the last two bytes of an FSAVE or an FSTENV operand are for any reason not writeable,
|
||||
or either of the last two bytes of an FRESTOR or FLDENV are for any reason not readable, the instruction
|
||||
is not restartable.
|
||||
|
||||
**Workaround**: This does not not affect typical systems with reasonably-assigned page access rights.
|
||||
In an obscure situation where this problem arises, a workaround is to avoid having the operand of these
|
||||
instructions span a page boundary. This can be accomplished by aligning these operands on any 128-byte boundary.
|
||||
|
|
@ -395,13 +391,11 @@ Here's what the world knew about 80386 problems in the B1 stepping, as of Decemb
|
|||
of maximum size (i.e. 0FFFFh for a 16-bit segment or 0FFFFFFFFh for a 32-bit segment) or within 108 bytes of
|
||||
maximum size, thus wrapping around to offset 0 of the segment. Since a wraparound situation is very abnormal
|
||||
for a compiler or programmer to create, this does not affect a typical system.
|
||||
|
||||
Formally, the 80386 architecture does not permit an operand (coprocessor operands included) to wrap around
|
||||
the end of a segment. If the user issues such an instruction nonetheless in a Protected Mode system, and
|
||||
the operand starts and ends in valid, present pages of a segment, BUT spans through an invalid or inaccessible
|
||||
page, the coprocessor may be put in an indeterminate state. In such cases, an FCLEX or FINIT instruction needs
|
||||
to be executed before any other coprocessor instruction is issued.
|
||||
|
||||
**Workaround**: In Real Mode, this is not a problem since protection is not enabled. In Protected Mode,
|
||||
this problem is avoided simply by not creating coprocessor operands which wrap around the end of the segment,
|
||||
or by aligning the base of all segments on page boundaries.
|
||||
|
|
@ -411,7 +405,6 @@ Here's what the world knew about 80386 problems in the B1 stepping, as of Decemb
|
|||
Furthermore, if the Double Fault entry in the IDT is a trap gate, a shutdown results. In a related topic,
|
||||
if the TSS Fault entry in the IDT is invalid for any reason (e.g. bad AR byte), then instead of a Double Fault
|
||||
(exception 8), a shutdown results.
|
||||
|
||||
**Workaround**: A working system, one that creates TSS segments of adequate size to hold the processor state
|
||||
(44 bytes for the TSS of a 16-bit task, 104 bytes for the TSS of a 32-bit task), will not encounter any problems
|
||||
here. A working system should also provide a valid gate (interrupt, trap, or task gate) in the IDT for exception 8.
|
||||
|
|
@ -422,37 +415,29 @@ Here's what the world knew about 80386 problems in the B1 stepping, as of Decemb
|
|||
next even-numbered iteration. If the REP MOVS ends with an odd number of iterations, and single-stepping or data
|
||||
breakpoints are enabled, then a single-step trap or data breakpoint trap on the final iteration will properly occur
|
||||
after the final, odd-numbered iteration.
|
||||
|
||||
**Workaround**: When using the Trap Flag or data breakpoints with a debugger utility, this minor variation of
|
||||
REP MOVS must be accepted, unless an effort is made to have the debugger emulate the REP MOVS rather than actually
|
||||
execute it.
|
||||
6. Task Switch to Virtual 8086 Mode Doesn't Update Prefetch Limit
|
||||
**Problem**: When a task switch to Virtual 8086 Mode is performed, the prefetch limit is not updated to become 0FFFFh,
|
||||
but instead remains at its previous value.
|
||||
|
||||
**Workaround**: Use the IRET instruction to transfer to Virtual 8086 Mode. Using IRET is the preferred method for
|
||||
most instances, especially when the master OS dispatches a Virtual 8086 Mode program, because IRET can cause the
|
||||
transition without a task switch.
|
||||
7. Wrong Register Size for String Instructions in Mixed 16/32-bit Addressing Systems
|
||||
**Problem**: If certain string and loop instructions are followed by instructions that either:
|
||||
|
||||
1) use a different address size (that is, if either the string instruction or the following instruction
|
||||
uses an address size prefix), or
|
||||
|
||||
2) reference the stack (e.g. PUSH/POP/CALL/RET) and the "B" bit in the SS descriptor is different from the address size used by the string
|
||||
instructions,
|
||||
|
||||
2) reference the stack (e.g. PUSH/POP/CALL/RET) and the "B" bit in the SS descriptor is different from the address
|
||||
size used by the string instructions,
|
||||
then one or more of [E]CX, [E]SI, or [E]DI is not updated properly. The size of the register (16 vs. 32) is
|
||||
taken from the following instruction rather than from the string or loop instruction. This could result in
|
||||
updating only the lower 16 bits of a 32-bit register, or in updating all 32 bits of a register being used as
|
||||
16 bits. The instructions (and registers) affected by this are:
|
||||
|
||||
MOVS ([E]DI), REP MOVS ([E]SI), STOS ([E]DI), INS ([E]DI), and REP INS ([E]CX).
|
||||
|
||||
**Workaround**: No workaround is necessary if all code is 16-bit or if all code is 32-bit. The problem only
|
||||
occurs if instructions with different address sizes are mixed together, or if a code segment of one size is used
|
||||
with a stack segment of the other size.
|
||||
|
||||
In a system which mixes address sizes, add a NOP after each of the above instructions and ensure that the NOP
|
||||
has the same address size as the string/loop (i.e., if the string/loop instruction includes an address prefix,
|
||||
place the same address prefix before the NOP; conversely, if the string/loop instruction does not have an address
|
||||
|
|
@ -464,19 +449,15 @@ Here's what the world knew about 80386 problems in the B1 stepping, as of Decemb
|
|||
old code segment. This can allow execution beyond the end of the new segment without triggering a segment limit
|
||||
violation. Or it can result in a spurious GP fault if the old and new segments overlap, and a prefetch occurs
|
||||
beyond the limit of the old segment.
|
||||
|
||||
Note that the prefetch limit is checked on the linear address, not by comparing IP to 0FFFFh.
|
||||
|
||||
**Workaround**: All existing 8086 programs use only 16-bit addressing, and thus will not execute code at offsets
|
||||
greater than 0FFFFh from the code segment base. Thus the lack of detection of walking off the end of a code segment
|
||||
should not impact working 8086 programs.
|
||||
|
||||
A workaround to the spurious GP fault, if it occurs, is to simply IRET back to the faulting instruction, since the
|
||||
IRET will correctly set the prefetch limit. If the fault handler has control of the single-step function, a very
|
||||
simple workaround is to attempt to single-step the faulting instruction. If the single-step succeeded, the handler
|
||||
could clear the fault, turn off single-stepping, and IRET. If a GP fault occurred attempting to single-step the
|
||||
instruction, a "real" GP fault is the cause.
|
||||
|
||||
If the fault handler cannot access the single-stepping function, it still can check for "real" GP faults which must
|
||||
be emulated by the master OS, for example, I/O instructions that need to be emulated, CLI/STI instructions that must
|
||||
be emulated, etc. If none of these faults are recognized, the fault handler can assume this errata caused the GP fault
|
||||
|
|
@ -484,7 +465,6 @@ Here's what the world knew about 80386 problems in the B1 stepping, as of Decemb
|
|||
9. Page Fault Error Code on Stack Not Reliable
|
||||
**Problem**: When a Page Fault (exception 14) occurs, the 3 defined bits in the error code may be unreliable
|
||||
if a certain sequence of prefetch is happening at the same time.
|
||||
|
||||
**Workaround**: Although the page fault error code pushed onto the page fault handler's stack can be unreliable,
|
||||
as described, the page fault linear address stored in register CR2 is always correct. The page fault handler should
|
||||
refer to the page fault linear address in CR2 to access the corresponding page table entry and thereby determine
|
||||
|
|
@ -494,23 +474,18 @@ Here's what the world knew about 80386 problems in the B1 stepping, as of Decemb
|
|||
or accessing coprocessor ports (I/O addresses 800000F8h-800000FFh) as a result of executing coprocessor opcodes,
|
||||
can generate incorrect I/O addresses if paging is enabled and the corresponding linear memory address is marked
|
||||
"present" and "dirty."
|
||||
|
||||
Furthermore, when paging has been enabled and is then turned off, paging translation continues to occur for memory
|
||||
or I/O cycles (I/O as described above) to linear addresses still stored in the TLB, but paging does not occur for
|
||||
linear addresses that result in a TLB miss.
|
||||
|
||||
**Workaround**: Unless paging is used, this item is not a problem. If paging is used but all I/O ports are below
|
||||
00001000h (as in a PC-DOS system), then I/O is no problem.
|
||||
|
||||
If paging is used and I/O ports exist in the range 0000l000h-0000FFFFh, then either have the memory pages at those
|
||||
linear addresses marked "not present" (to avoid having those pages table entries cached in the TLB), or if "present,"
|
||||
have those pages mapped such that bits 12-15 of the physical address equal bits 12-15 of the linear address.
|
||||
Alternatively, re-assign any I/O ports in the range 00001000h-0000FFFFh to below 00001000h.
|
||||
|
||||
If paging is used and the coprocessor is also used, then have the memory page at linear address 80000xxxh either
|
||||
marked "not present" (to avoid having that page table entry cached in the TLB), or if "present," have the page
|
||||
mapped such that bit 31 (the most significant bit) of that page's physical address is a 1.
|
||||
|
||||
To completely disable 80386 paging when paging was previously enabled, the 80386 TLB should be flushed immediately
|
||||
after resetting the~PG bit in CRO. The TLB can be flushed, you recall, by writing a Page Table Directory base address
|
||||
to register CR3.
|
||||
|
|
@ -521,7 +496,6 @@ Here's what the world knew about 80386 problems in the B1 stepping, as of Decemb
|
|||
But in the case of a REP INS instruction, ECX is not updated correctly and is 0FFFFFFFFh (or CX is 0FFFFh in case of
|
||||
16-bit operations). It should be noted that the REP INS executes the correct number of iterations and EDI (or DI)
|
||||
is updated properly.
|
||||
|
||||
**Workaround**: After a REP INS instruction, do not rely on ECX (or CX) being zero. Hence, a new count (if any)
|
||||
should be MOVed into ECX, rather than being ADDed into ECX.
|
||||
12. NMI Doesn't Always Bring Chip Out of Shutdown in Obscure Condition with Paging Enabled
|
||||
|
|
@ -530,51 +504,37 @@ Here's what the world knew about 80386 problems in the B1 stepping, as of Decemb
|
|||
and a TLB miss occurs when accessing the null descriptor slot, the chip enters shutdown as it should in this case.
|
||||
In this specific case however, an incoming NMI will not be able to bring the 386 out of shutdown. In this specific
|
||||
case, only reset will bring the 386 out of shutdown.
|
||||
|
||||
**Workaround**: Ensure that the IDT gate for the Double Fault Handler has a non-null selectors for CS, and that
|
||||
SS of the destination level is also non-null.
|
||||
13. HOLD Input During Protected Mode Interlevel IRET when Paging is Enabled
|
||||
**Problem**: Under specific situations involving paging and the page privilege bits, the HOLD input, and a RET
|
||||
or IRET instruction performing an inter-level return to level 3, a problem can develop. These situations can be
|
||||
avoided by the workarounds given.
|
||||
|
||||
The first situation, when the inner level stack (levels 0, 1, and 2) is not dword aligned (or not word aligned
|
||||
in the case of a 16-bit [I]RET), requires that several conditions occur simultaneously:
|
||||
|
||||
1) Paging must be enabled, and the page table and directory entries for the inner level stacks must be marked
|
||||
as supervisor access only.
|
||||
|
||||
2) The software must execute an inter-level RET or IRET to a Protected Mode program at privilege level 3.
|
||||
An inter-level IRET to Virtual 8086 Mode does not exhibit this problem. An inter-level RET or IRET to level 1
|
||||
or 2 does not exhibit this problem.
|
||||
|
||||
3) The inner level stack must be unaligned to a dword boundary (word boundary for a 16-bit [I]RET).
|
||||
|
||||
When the first situation occurs, a page fault (exception 14) occurs spuriously, indicating a page level
|
||||
protection violation during a "user" level read of the inner level stack.
|
||||
|
||||
The second situation, whether or not the inner level stack is dword aligned (or word aligned in the case of a
|
||||
16-bit [I]RET), also requires that several conditions occur simultaneously:
|
||||
|
||||
1) Paging must be enabled, and the page table and directory entries for the inner level stacks must be marked
|
||||
as supervisor access only.
|
||||
|
||||
2) The software must execute an inter-level RET or IRET to a Protected Mode program at privilege level 3.
|
||||
An inter-level IRET to Virtual 8086 Mode does not exhibit this problem. An inter-level RET or IRET to level 1
|
||||
or 2 does not exhibit this problem.
|
||||
|
||||
3) The bus HOLD input must be asserted during the read, cycle which pops ESP (or SP) off the inner stack as a
|
||||
result of a RET or IRET instruction.
|
||||
|
||||
When the second situation occurs, no exception is generated, but the processor will drive an incorrect physical
|
||||
address during the read cycle in which SS is popped from the inner level stack.
|
||||
|
||||
**Workarounds**: A software workaround to both situations is to mark all pages which contain the inner level
|
||||
stacks as user readable. This prevents either the first or second situation from occurring. The segmentation
|
||||
system can be used to prevent user access to the linear addresses containing the inner-level stacks.
|
||||
|
||||
A workaround if not using the HOLD input is merely to keep the inner-level stacks aligned.
|
||||
|
||||
A Hardware workaround if using the HOLD input but not using the software workaround above is the following:
|
||||
Since the problem occurs during the first cycle after a locked cycle to read the CS descriptor, a hardware
|
||||
workaround is to prevent a HOLD request from hitting the processor during bus cycle following a LOCKed cycle.
|
||||
|
|
@ -588,7 +548,6 @@ Here's what the world knew about 80386 problems in the B1 stepping, as of Decemb
|
|||
perform a stack operation, such as PUSH or POP (see exact list below), the value of the [E]SP register may be
|
||||
incorrect after the stack operation. Note that stack operations resulting from interrupts or exceptions following
|
||||
LSL do update [E]SP correctly.
|
||||
|
||||
**Workaround**: Do not immediately follow the Protected Mode LSL instruction with any of the following stack
|
||||
operation instructions: IRET (intra-task), POPA, POPF, POP (mem, reg, seg-reg), RET (intrasegment or intersegment),
|
||||
CALL (direct intrasegment, direct intersegment, indirect intrasegment via reg), ENTER, PUSHA, PUSHF, PUSH (mem,
|
||||
|
|
@ -600,7 +559,6 @@ Here's what the world knew about 80386 problems in the B1 stepping, as of Decemb
|
|||
**Problem**: The Protected Mode instructions LSL, LAR, VERR or VERW executed with a null selector (i.e. bits
|
||||
15 through 2 of the selector set to zero) as the operand will operate on the descriptor at entry 0 of the GDT
|
||||
instead of unconditionally clearing the ZF flag.
|
||||
|
||||
**Workaround**: The "null descriptor" (i.e. the descriptor at entry 0 of the GDT) should be initialized to all
|
||||
zeroes. If the "null descriptor" is initialized to all zeroes (i.e. an invalid value), the access made by these
|
||||
instructions to the "null descriptor" will fail (since these instructions only operate on valid descriptors).
|
||||
|
|
@ -610,18 +568,15 @@ Here's what the world knew about 80386 problems in the B1 stepping, as of Decemb
|
|||
16. "Not Present" LDT in VM86 Task Raises Wrong Exception
|
||||
**Problem**: A task switch to a VM86 task that has a "not present" LDT descriptor will cause a Segment Not Present
|
||||
fault (exception 11) rather than an Invalid TSS fault (exception 10).
|
||||
|
||||
**Workaround**: The simplest workaround is to use a NULL selector for the LDT in a VM86 task, since the LDT is
|
||||
not used when executing in Virtual 86 mode. However, if an interrupt or exception occurs, the processor will switch
|
||||
out of Virtual 86 mode, into protected mode to handle the interrupt, without switching tasks. Thus, the operating
|
||||
system should be structured so that all Interrupt and Trap gates active when executing a VM86 task reference segments
|
||||
in the GDT.
|
||||
|
||||
If an LDT must be supplied for a task that executes in Virtual 86 mode, there are several easy workarounds. One
|
||||
is to ensure that LDT segments are never marked "not present" in their segment descriptors. Paging is not affected
|
||||
by this errata. LDT segments can be paged out and marked "not present" in their page descriptors in systems which
|
||||
use paging.
|
||||
|
||||
If the operating system must mark the LDT segment descriptor "not present", the "not present" (exception 11)
|
||||
handler must be able to handle the case of a "not present" LDT during a task switch. The "not present" exception
|
||||
is reported with the LDT selector as the error code and with the VM bit set to 1 in the EFLAGS image of the caller.
|
||||
|
|
@ -632,20 +587,14 @@ Here's what the world knew about 80386 problems in the B1 stepping, as of Decemb
|
|||
and the second byte is located on a page or segment which would create a fault, then the processor will hang when
|
||||
it tries to signal the fault. The processor remains stopped until an interrupt, NMI, or RESET occurs. This errata
|
||||
applies only to coprocessor instructions in systems which use virtual memory.
|
||||
|
||||
**Workaround**: In virtual memory systems, the time-slice or watchdog timer provides an easy workaround, since a
|
||||
timer interrupt will always cause the processor to begin interrupt processing. The timer routine should test the
|
||||
following conditions to determine if this errata was encountered.
|
||||
|
||||
1) The saved CS:EIP must point within 8 bytes of the end of a page.
|
||||
|
||||
2) The last byte within the page must contain an ESC opcode.
|
||||
|
||||
3) All bytes between the saved CS:EIP and the ESC opcode must contain valid prefix opcodes (segment override 26h,
|
||||
2Eh, 36h, 3Eh, 64h, 65h, address size override 67h, operand size override 66h).
|
||||
|
||||
4) The next page is not present, or not accessable.
|
||||
|
||||
If all four conditions are true, then the timer routine can assume this errata was encountered, and signal a page
|
||||
fault, which will clear the condition. This workaround should be placed in the Operating System, so that applications
|
||||
programs are unaffected.
|
||||
|
|
@ -653,7 +602,6 @@ Here's what the world knew about 80386 problems in the B1 stepping, as of Decemb
|
|||
**Problem**: If a second page fault occurs, while the processor is attempting to enter the service routine for the
|
||||
first, then the processor will invoke the page fault (exception 14) handler a second time, rather than the double
|
||||
fault (exception 8) handler. A subsequent fault, though, will lead to shutdown.
|
||||
|
||||
**Workaround**: No workaround is necessary in a working system.
|
||||
|
||||
An errata update dated March 26, 1987, produced internally by IBM rather than Intel, noted two additional
|
||||
|
|
|
|||
Loading…
Reference in a new issue