80386 B1 errata details
This commit is contained in:
parent
8112026eba
commit
169a1d72ad
1 changed files with 276 additions and 8 deletions
|
|
@ -4,11 +4,12 @@ Assembling a detailed and accurate history of the 80386, including a complete li
|
|||
(revisions), when they were released, what "errata" (problems) each stepping suffered from, and which of those
|
||||
problems were fixed by a later stepping, seems virtually impossible at this late date.
|
||||
|
||||
I won't make the attempt here, either. I'm just going to throw together everything I know into one pile,
|
||||
as I begin adding 80386 support to PCjs.
|
||||
I won't make the attempt here, either. Using information from various sources, I'll start with an overview
|
||||
of the steppings (revision levels), including how each stepping was externally marked and internally identified,
|
||||
along with lists of associated errata, then move on to more detailed errata information (from Intel's own documents),
|
||||
and end with a summary of undocumented 80386 instructions.
|
||||
|
||||
Let's start with steppings (revision levels), then move on to errata, and finally undocumented and/or deprecated
|
||||
instructions.
|
||||
Until further notice, this document is a work-in-progress.
|
||||
|
||||
### Steppings
|
||||
|
||||
|
|
@ -363,30 +364,297 @@ identifier 0x05.
|
|||
|
||||
### Errata
|
||||
|
||||
It's fair to say that the 80386 B1 stepping had a lot of problems. Although, for a CPU of significantly
|
||||
greater complexity than its predecessors, that doesn't seem terribly surprising.
|
||||
It's fair to say that early 80386 steppings had a lot of problems. But that's not terribly surprising,
|
||||
considering the 80386's leap in functionality: a new 32-bit architecture that included paging and virtual
|
||||
8086 mode, while also maintaining 16-bit protected-mode compatibility with the 80286, not to mention
|
||||
real-mode compatibility with the 8086.
|
||||
|
||||
From the aforementioned December 17, 1986 document, here's what the world knew about 80386-B1 problems at
|
||||
that time:
|
||||
Here's what the world knew about 80386 problems in the B1 stepping, as of December 17, 1986:
|
||||
|
||||
1. Opcode Field Incorrect for FSAVE and FSTENV
|
||||
**Problem**: If an FSAVE or an FSTENV is executed in REAL mode or in VIRTUAL 8086 mode, the opcode field
|
||||
stored in memory is incorrect if it should have referred to a coprocessor instruction which transfers
|
||||
either two bytes or ten bytes from memory to the coprocessor. The instruction and operand linear address
|
||||
fields are correctly stored. Note that coprocessor error-handling routines are the only routines possibly
|
||||
affected. Also note that the problem does not occur in PROTECTED mode programs (since no opcode is saved
|
||||
by FSAVE or FSTENV in that case).
|
||||
|
||||
**Workaround**: In REAL mode or in VIRTUAL 8086 mode, the instruction linear address field can be used to
|
||||
read the opcode from memory. Note that the two bytes fetched need to be swapped to yield the image that
|
||||
FSAVE and FSTENV normally stores.
|
||||
2. FSAVE, FRESTOR, FSTENV and FLDENV Anomolies [sic] with Paging
|
||||
**Problem**: If either of the last two bytes of an FSAVE or an FSTENV operand are for any reason not writeable,
|
||||
or either of the last two bytes of an FRESTOR or FLDENV are for any reason not readable, the instruction
|
||||
is not restartable.
|
||||
|
||||
**Workaround**: This does not not affect typical systems with reasonably-assigned page access rights.
|
||||
In an obscure situation where this problem arises, a workaround is to avoid having the operand of these
|
||||
instructions span a page boundary. This can be accomplished by aligning these operands on any 128-byte boundary.
|
||||
3. Wraparound Coprocessor Operands
|
||||
**Problem**: This can affect only situations where a coprocessor operand straddles the limit of a segment
|
||||
of maximum size (i.e. 0FFFFh for a 16-bit segment or 0FFFFFFFFh for a 32-bit segment) or within 108 bytes of
|
||||
maximum size, thus wrapping around to offset 0 of the segment. Since a wraparound situation is very abnormal
|
||||
for a compiler or programmer to create, this does not affect a typical system.
|
||||
|
||||
Formally, the 80386 architecture does not permit an operand (coprocessor operands included) to wrap around
|
||||
the end of a segment. If the user issues such an instruction nonetheless in a Protected Mode system, and
|
||||
the operand starts and ends in valid, present pages of a segment, BUT spans through an invalid or inaccessible
|
||||
page, the coprocessor may be put in an indeterminate state. In such cases, an FCLEX or FINIT instruction needs
|
||||
to be executed before any other coprocessor instruction is issued.
|
||||
|
||||
**Workaround**: In Real Mode, this is not a problem since protection is not enabled. In Protected Mode,
|
||||
this problem is avoided simply by not creating coprocessor operands which wrap around the end of the segment,
|
||||
or by aligning the base of all segments on page boundaries.
|
||||
4. IRET to TSS with Limit too Small
|
||||
**Problem**: If an IRET performs a task switch to a TSS of proper descriptor type but invalid (too small) limit,
|
||||
a Double Fault (exception 8) will result instead of a Invalid TSS Fault (exception 10) as should result.
|
||||
Furthermore, if the Double Fault entry in the IDT is a trap gate, a shutdown results. In a related topic,
|
||||
if the TSS Fault entry in the IDT is invalid for any reason (e.g. bad AR byte), then instead of a Double Fault
|
||||
(exception 8), a shutdown results.
|
||||
|
||||
**Workaround**: A working system, one that creates TSS segments of adequate size to hold the processor state
|
||||
(44 bytes for the TSS of a 16-bit task, 104 bytes for the TSS of a 32-bit task), will not encounter any problems
|
||||
here. A working system should also provide a valid gate (interrupt, trap, or task gate) in the IDT for exception 8.
|
||||
5. Single-Stepping First Iteration of REP MOVS
|
||||
**Problem**: If a REPeated MOVS instruction is executed when single-stepping is enabled (TF = 1 in EFLAGS register),
|
||||
a single-step trap (exception 1) is taken every two move steps, but should occur each move step. Also, if a data
|
||||
breakpoint is hit during a odd iteration number of REP MOVS, the data breakpoint trap is not taken until after the
|
||||
next even-numbered iteration. If the REP MOVS ends with an odd number of iterations, and single-stepping or data
|
||||
breakpoints are enabled, then a single-step trap or data breakpoint trap on the final iteration will properly occur
|
||||
after the final, odd-numbered iteration.
|
||||
|
||||
**Workaround**: When using the Trap Flag or data breakpoints with a debugger utility, this minor variation of
|
||||
REP MOVS must be accepted, unless an effort is made to have the debugger emulate the REP MOVS rather than actually
|
||||
execute it.
|
||||
6. Task Switch to Virtual 8086 Mode Doesn't Update Prefetch Limit
|
||||
**Problem**: When a task switch to Virtual 8086 Mode is performed, the prefetch limit is not updated to become 0FFFFh,
|
||||
but instead remains at its previous value.
|
||||
|
||||
**Workaround**: Use the IRET instruction to transfer to Virtual 8086 Mode. Using IRET is the preferred method for
|
||||
most instances, especially when the master OS dispatches a Virtual 8086 Mode program, because IRET can cause the
|
||||
transition without a task switch.
|
||||
7. Wrong Register Size for String Instructions in Mixed 16/32-bit Addressing Systems
|
||||
**Problem**: If certain string and loop instructions are followed by instructions that either:
|
||||
|
||||
1) use a different address size (that is, if either the string instruction or the following instruction
|
||||
uses an address size prefix), or
|
||||
|
||||
2) reference the stack (e.g. PUSH/POP/CALL/RET) and the "B" bit in the SS descriptor is different from the address size used by the string
|
||||
instructions,
|
||||
|
||||
then one or more of [E]CX, [E]SI, or [E]DI is not updated properly. The size of the register (16 vs. 32) is
|
||||
taken from the following instruction rather than from the string or loop instruction. This could result in
|
||||
updating only the lower 16 bits of a 32-bit register, or in updating all 32 bits of a register being used as
|
||||
16 bits. The instructions (and registers) affected by this are:
|
||||
|
||||
MOVS ([E]DI), REP MOVS ([E]SI), STOS ([E]DI), INS ([E]DI), and REP INS ([E]CX).
|
||||
|
||||
**Workaround**: No workaround is necessary if all code is 16-bit or if all code is 32-bit. The problem only
|
||||
occurs if instructions with different address sizes are mixed together, or if a code segment of one size is used
|
||||
with a stack segment of the other size.
|
||||
|
||||
In a system which mixes address sizes, add a NOP after each of the above instructions and ensure that the NOP
|
||||
has the same address size as the string/loop (i.e., if the string/loop instruction includes an address prefix,
|
||||
place the same address prefix before the NOP; conversely, if the string/loop instruction does not have an address
|
||||
prefix, do not place a prefix before the NOP).
|
||||
8. FAR Jump Located Near Page Boundary in Virtual 8086 Mode Paged Systems
|
||||
**Problem**: In Virtual 8086 Mode, if a direct FAR jump (opcode EAh) instruction is located at the end of a
|
||||
page (or within 16 bytes of the end), and the next page is not cached in the TLB, the prefetcher limit is not
|
||||
set by the FAR jump instruction to the "end" on the new code segment, but rather is left at the "end" of the
|
||||
old code segment. This can allow execution beyond the end of the new segment without triggering a segment limit
|
||||
violation. Or it can result in a spurious GP fault if the old and new segments overlap, and a prefetch occurs
|
||||
beyond the limit of the old segment.
|
||||
|
||||
Note that the prefetch limit is checked on the linear address, not by comparing IP to 0FFFFh.
|
||||
|
||||
**Workaround**: All existing 8086 programs use only 16-bit addressing, and thus will not execute code at offsets
|
||||
greater than 0FFFFh from the code segment base. Thus the lack of detection of walking off the end of a code segment
|
||||
should not impact working 8086 programs.
|
||||
|
||||
A workaround to the spurious GP fault, if it occurs, is to simply IRET back to the faulting instruction, since the
|
||||
IRET will correctly set the prefetch limit. If the fault handler has control of the single-step function, a very
|
||||
simple workaround is to attempt to single-step the faulting instruction. If the single-step succeeded, the handler
|
||||
could clear the fault, turn off single-stepping, and IRET. If a GP fault occurred attempting to single-step the
|
||||
instruction, a "real" GP fault is the cause.
|
||||
|
||||
If the fault handler cannot access the single-stepping function, it still can check for "real" GP faults which must
|
||||
be emulated by the master OS, for example, I/O instructions that need to be emulated, CLI/STI instructions that must
|
||||
be emulated, etc. If none of these faults are recognized, the fault handler can assume this errata caused the GP fault
|
||||
and simply IRET back to the instruction.
|
||||
9. Page Fault Error Code on Stack Not Reliable
|
||||
**Problem**: When a Page Fault (exception 14) occurs, the 3 defined bits in the error code may be unreliable
|
||||
if a certain sequence of prefetch is happening at the same time.
|
||||
|
||||
**Workaround**: Although the page fault error code pushed onto the page fault handler's stack can be unreliable,
|
||||
as described, the page fault linear address stored in register CR2 is always correct. The page fault handler should
|
||||
refer to the page fault linear address in CR2 to access the corresponding page table entry and thereby determine
|
||||
whether the page fault was due to a page "not present" condition, or to a usage violation.
|
||||
10. Certain I/O Addresses Incorrect when Paging is Enabled
|
||||
**Problem**: When Paging is enabled, accessing I/O addresses in the range 00001000h-0000FFFFh (4K through 64K-1)
|
||||
or accessing coprocessor ports (I/O addresses 800000F8h-800000FFh) as a result of executing coprocessor opcodes,
|
||||
can generate incorrect I/O addresses if paging is enabled and the corresponding linear memory address is marked
|
||||
"present" and "dirty."
|
||||
|
||||
Furthermore, when paging has been enabled and is then turned off, paging translation continues to occur for memory
|
||||
or I/O cycles (I/O as described above) to linear addresses still stored in the TLB, but paging does not occur for
|
||||
linear addresses that result in a TLB miss.
|
||||
|
||||
**Workaround**: Unless paging is used, this item is not a problem. If paging is used but all I/O ports are below
|
||||
00001000h (as in a PC-DOS system), then I/O is no problem.
|
||||
|
||||
If paging is used and I/O ports exist in the range 0000l000h-0000FFFFh, then either have the memory pages at those
|
||||
linear addresses marked "not present" (to avoid having those pages table entries cached in the TLB), or if "present,"
|
||||
have those pages mapped such that bits 12-15 of the physical address equal bits 12-15 of the linear address.
|
||||
Alternatively, re-assign any I/O ports in the range 00001000h-0000FFFFh to below 00001000h.
|
||||
|
||||
If paging is used and the coprocessor is also used, then have the memory page at linear address 80000xxxh either
|
||||
marked "not present" (to avoid having that page table entry cached in the TLB), or if "present," have the page
|
||||
mapped such that bit 31 (the most significant bit) of that page's physical address is a 1.
|
||||
|
||||
To completely disable 80386 paging when paging was previously enabled, the 80386 TLB should be flushed immediately
|
||||
after resetting the~PG bit in CRO. The TLB can be flushed, you recall, by writing a Page Table Directory base address
|
||||
to register CR3.
|
||||
11. Wrong ECX Update by REP INS
|
||||
**Problem**: The ECX register (or CX in case of 16-bit operations) is not updated properly in the case of a
|
||||
REP INS instruction (INPut string instruction with any REPeat prefix) that is followed by an early-start instruction
|
||||
(e.g. PUSH, POP or memory reference instructions). After any REP-prefixed instruction, ECX is supposed to be 0 (null).
|
||||
But in the case of a REP INS instruction, ECX is not updated correctly and is 0FFFFFFFFh (or CX is 0FFFFh in case of
|
||||
16-bit operations). It should be noted that the REP INS executes the correct number of iterations and EDI (or DI)
|
||||
is updated properly.
|
||||
|
||||
**Workaround**: After a REP INS instruction, do not rely on ECX (or CX) being zero. Hence, a new count (if any)
|
||||
should be MOVed into ECX, rather than being ADDed into ECX.
|
||||
12. NMI Doesn't Always Bring Chip Out of Shutdown in Obscure Condition with Paging Enabled
|
||||
**Problem**: If paging is enabled, and if the IDT gate for the Double Fault handler (the gate for exception 8)
|
||||
points to the null descriptor slot, descriptor 0, in the GDT (this would be very a strange way to set up a system),
|
||||
and a TLB miss occurs when accessing the null descriptor slot, the chip enters shutdown as it should in this case.
|
||||
In this specific case however, an incoming NMI will not be able to bring the 386 out of shutdown. In this specific
|
||||
case, only reset will bring the 386 out of shutdown.
|
||||
|
||||
**Workaround**: Ensure that the IDT gate for the Double Fault Handler has a non-null selectors for CS, and that
|
||||
SS of the destination level is also non-null.
|
||||
13. HOLD Input During Protected Mode Interlevel IRET when Paging is Enabled
|
||||
**Problem**: Under specific situations involving paging and the page privilege bits, the HOLD input, and a RET
|
||||
or IRET instruction performing an inter-level return to level 3, a problem can develop. These situations can be
|
||||
avoided by the workarounds given.
|
||||
|
||||
The first situation, when the inner level stack (levels 0, 1, and 2) is not dword aligned (or not word aligned
|
||||
in the case of a 16-bit [I]RET), requires that several conditions occur simultaneously:
|
||||
|
||||
1) Paging must be enabled, and the page table and directory entries for the inner level stacks must be marked
|
||||
as supervisor access only.
|
||||
|
||||
2) The software must execute an inter-level RET or IRET to a Protected Mode program at privilege level 3.
|
||||
An inter-level IRET to Virtual 8086 Mode does not exhibit this problem. An inter-level RET or IRET to level 1
|
||||
or 2 does not exhibit this problem.
|
||||
|
||||
3) The inner level stack must be unaligned to a dword boundary (word boundary for a 16-bit [I]RET).
|
||||
|
||||
When the first situation occurs, a page fault (exception 14) occurs spuriously, indicating a page level
|
||||
protection violation during a "user" level read of the inner level stack.
|
||||
|
||||
The second situation, whether or not the inner level stack is dword aligned (or word aligned in the case of a
|
||||
16-bit [I]RET), also requires that several conditions occur simultaneously:
|
||||
|
||||
1) Paging must be enabled, and the page table and directory entries for the inner level stacks must be marked
|
||||
as supervisor access only.
|
||||
|
||||
2) The software must execute an inter-level RET or IRET to a Protected Mode program at privilege level 3.
|
||||
An inter-level IRET to Virtual 8086 Mode does not exhibit this problem. An inter-level RET or IRET to level 1
|
||||
or 2 does not exhibit this problem.
|
||||
|
||||
3) The bus HOLD input must be asserted during the read, cycle which pops ESP (or SP) off the inner stack as a
|
||||
result of a RET or IRET instruction.
|
||||
|
||||
When the second situation occurs, no exception is generated, but the processor will drive an incorrect physical
|
||||
address during the read cycle in which SS is popped from the inner level stack.
|
||||
|
||||
**Workarounds**: A software workaround to both situations is to mark all pages which contain the inner level
|
||||
stacks as user readable. This prevents either the first or second situation from occurring. The segmentation
|
||||
system can be used to prevent user access to the linear addresses containing the inner-level stacks.
|
||||
|
||||
A workaround if not using the HOLD input is merely to keep the inner-level stacks aligned.
|
||||
|
||||
A Hardware workaround if using the HOLD input but not using the software workaround above is the following:
|
||||
Since the problem occurs during the first cycle after a locked cycle to read the CS descriptor, a hardware
|
||||
workaround is to prevent a HOLD request from hitting the processor during bus cycle following a LOCKed cycle.
|
||||
This can be accomplished with a latch that delays the LOCK# signal through a flip-flop clocked by READY# to
|
||||
gate a HOLD request going into the chip. This will prevent a hold request from getting to the 80386 until after
|
||||
the completion of the first cycle after a LOCKed cycle. For the hardware workaround to be sufficient, all stacks
|
||||
must be properly aligned, and BS16# must be tied inactive.
|
||||
14. Protected Mode LSL Instruction Should not be Followed by PUSH/POP
|
||||
**Problem**: This item pertains only to Protected Mode. If the Protected Mode LSL instruction (Load Segment
|
||||
Limit instruction, executable only in Protected Mode) is immediately followed by certain instructions that
|
||||
perform a stack operation, such as PUSH or POP (see exact list below), the value of the [E]SP register may be
|
||||
incorrect after the stack operation. Note that stack operations resulting from interrupts or exceptions following
|
||||
LSL do update [E]SP correctly.
|
||||
|
||||
**Workaround**: Do not immediately follow the Protected Mode LSL instruction with any of the following stack
|
||||
operation instructions: IRET (intra-task), POPA, POPF, POP (mem, reg, seg-reg), RET (intrasegment or intersegment),
|
||||
CALL (direct intrasegment, direct intersegment, indirect intrasegment via reg), ENTER, PUSHA, PUSHF, PUSH (mem,
|
||||
reg, seg-reg, immed). Other instructions that operate on the stack (e.g. CALL indirect via memory, and LEAVE) can
|
||||
be used safely after the Protected Mode LSL. Note that even if a forbidden instruction immediately follows LSL,
|
||||
[E]SP may still be updated correctly, since this problem is data-dependent and only occurs if the LSL operation
|
||||
succeeded (i.e. if LSL set the ZF flag).
|
||||
15. LSL/LAR/VERR/VERW Instructions Malfunction with Null Selector
|
||||
**Problem**: The Protected Mode instructions LSL, LAR, VERR or VERW executed with a null selector (i.e. bits
|
||||
15 through 2 of the selector set to zero) as the operand will operate on the descriptor at entry 0 of the GDT
|
||||
instead of unconditionally clearing the ZF flag.
|
||||
|
||||
**Workaround**: The "null descriptor" (i.e. the descriptor at entry 0 of the GDT) should be initialized to all
|
||||
zeroes. If the "null descriptor" is initialized to all zeroes (i.e. an invalid value), the access made by these
|
||||
instructions to the "null descriptor" will fail (since these instructions only operate on valid descriptors).
|
||||
The failure will be reported with ZF cleared, which is the desired behavior when the operand is a null selector.
|
||||
Note that many systems already have the "null descriptor" in the GDT initialized to zeroes, as is desired for
|
||||
this workaround.
|
||||
16. "Not Present" LDT in VM86 Task Raises Wrong Exception
|
||||
**Problem**: A task switch to a VM86 task that has a "not present" LDT descriptor will cause a Segment Not Present
|
||||
fault (exception 11) rather than an Invalid TSS fault (exception 10).
|
||||
|
||||
**Workaround**: The simplest workaround is to use a NULL selector for the LDT in a VM86 task, since the LDT is
|
||||
not used when executing in Virtual 86 mode. However, if an interrupt or exception occurs, the processor will switch
|
||||
out of Virtual 86 mode, into protected mode to handle the interrupt, without switching tasks. Thus, the operating
|
||||
system should be structured so that all Interrupt and Trap gates active when executing a VM86 task reference segments
|
||||
in the GDT.
|
||||
|
||||
If an LDT must be supplied for a task that executes in Virtual 86 mode, there are several easy workarounds. One
|
||||
is to ensure that LDT segments are never marked "not present" in their segment descriptors. Paging is not affected
|
||||
by this errata. LDT segments can be paged out and marked "not present" in their page descriptors in systems which
|
||||
use paging.
|
||||
|
||||
If the operating system must mark the LDT segment descriptor "not present", the "not present" (exception 11)
|
||||
handler must be able to handle the case of a "not present" LDT during a task switch. The "not present" exception
|
||||
is reported with the LDT selector as the error code and with the VM bit set to 1 in the EFLAGS image of the caller.
|
||||
Since a VM86 task cannot normally raise a "not present" fault, the "not present" exception handler can detect this
|
||||
case by checking if the stored VM bit is set. If so, the fault can be redirected to the TSS Fault handler.
|
||||
17. Coprocessor Instructions Crossing Page/Segment Boundaries
|
||||
**Problem**: If the first byte of a coprocessor (ESC) instruction is located on the last byte of a page or segment,
|
||||
and the second byte is located on a page or segment which would create a fault, then the processor will hang when
|
||||
it tries to signal the fault. The processor remains stopped until an interrupt, NMI, or RESET occurs. This errata
|
||||
applies only to coprocessor instructions in systems which use virtual memory.
|
||||
|
||||
**Workaround**: In virtual memory systems, the time-slice or watchdog timer provides an easy workaround, since a
|
||||
timer interrupt will always cause the processor to begin interrupt processing. The timer routine should test the
|
||||
following conditions to determine if this errata was encountered.
|
||||
|
||||
1) The saved CS:EIP must point within 8 bytes of the end of a page.
|
||||
|
||||
2) The last byte within the page must contain an ESC opcode.
|
||||
|
||||
3) All bytes between the saved CS:EIP and the ESC opcode must contain valid prefix opcodes (segment override 26h,
|
||||
2Eh, 36h, 3Eh, 64h, 65h, address size override 67h, operand size override 66h).
|
||||
|
||||
4) The next page is not present, or not accessable.
|
||||
|
||||
If all four conditions are true, then the timer routine can assume this errata was encountered, and signal a page
|
||||
fault, which will clear the condition. This workaround should be placed in the Operating System, so that applications
|
||||
programs are unaffected.
|
||||
18. Double Page Faults Do Not Raise Double Fault Exception
|
||||
**Problem**: If a second page fault occurs, while the processor is attempting to enter the service routine for the
|
||||
first, then the processor will invoke the page fault (exception 14) handler a second time, rather than the double
|
||||
fault (exception 8) handler. A subsequent fault, though, will lead to shutdown.
|
||||
|
||||
**Workaround**: No workaround is necessary in a working system.
|
||||
|
||||
An errata update dated March 26, 1987, produced internally by IBM rather than Intel, noted two additional
|
||||
issues:
|
||||
|
|
|
|||
Loading…
Reference in a new issue