Deduped images

This commit is contained in:
James Gregory 2013-12-30 12:25:34 +11:00
commit 9e52de9586
233 changed files with 128 additions and 128 deletions

View file

@ -53,7 +53,7 @@
<P>Notice that the above definition most emphatically does <I>not</I> say anything about making the software as fast as possible. It also does not say anything about using assembly language, or an optimizing compiler, or, for that matter, a compiler at all. It also doesn&#146;t say anything about how the code was designed and written. What it does say is that high-performance code shouldn&#146;t get in the user&#146;s way&#151;and that&#146;s <I>all</I>.</P>
<P>That&#146;s an important distinction, because all too many programmers think that assembly language, or the right compiler, or a particular high-level language, or a certain design approach is the answer to creating high-performance code. They&#146;re not, any more than choosing a certain set of tools is the key to building a house. You do indeed need tools to build a house, but any of many sets of tools will do. You also need a blueprint, an understanding of everything that goes into a house, and the ability to <I>use</I> the tools.</P>
<P>Likewise, high-performance programming requires a clear understanding of the purpose of the software being built, an overall program design, algorithms for implementing particular tasks, an understanding of what the computer can do and of what all relevant software is doing&#151;<I>and</I> solid programming skills, preferably using an optimizing compiler or assembly language. The optimization at the end is just the finishing touch, however.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/01-01i.jpg"><TD WIDTH="95%"><SMALL><I>Without good design, good algorithms, and complete understanding of the program&#146;s operation, your carefully optimized code will amount to one of mankind&#146;s least fruitful creations&#151;a fast slow program</I>.</SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>Without good design, good algorithms, and complete understanding of the program&#146;s operation, your carefully optimized code will amount to one of mankind&#146;s least fruitful creations&#151;a fast slow program</I>.</SMALL>
</TABLE>
<P>&#147;What&#146;s a fast slow program?&#148; you ask. That&#146;s a good question, and a brief (true) story is perhaps the best answer.
</P>

View file

@ -109,7 +109,7 @@ main(int argc, char *argv[]) {
</PRE>
<!-- END CODE //-->
<P>Table 1.1 shows the time taken for Listing 1.1 to generate a checksum of the WordPerfect version 4.2 thesaurus file, TH.WP (362,293 bytes in size), on a 10 MHz AT machine of no special parentage. Execution times are given for Listing 1.1 compiled with Borland and Microsoft compilers, with optimization both on and off; all four times are pretty much the same, however, and all are much too slow to be acceptable. Listing 1.1 requires over two and one-half minutes to checksum <I>one</I> file!</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/01-02i.jpg"><TD WIDTH="95%"><SMALL><I>Listings 1.2 and 1.3 form the C/assembly equivalent to Listing 1.1, and Listings 1.6 and 1.7 form the C/assembly equivalent to Listing 1.5.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>Listings 1.2 and 1.3 form the C/assembly equivalent to Listing 1.1, and Listings 1.6 and 1.7 form the C/assembly equivalent to Listing 1.5.</I></SMALL>
</TABLE>
<P>These results make it clear that it&#146;s folly to rely on your compiler&#146;s optimization to make your programs fast. Listing 1.1 is simply poorly designed, and no amount of compiler optimization will compensate for that failing. To drive home the point, conListings 1.2 and 1.3, which together are equivalent to Listing 1.1 except that the entire checksum loop is written in tight assembly code. The assembly language implementation is indeed faster than any of the C versions, as shown in Table 1.1, but it&#146;s less than 10 percent faster, and it&#146;s still unacceptably slow.
</P><P><BR></P>

View file

@ -37,7 +37,7 @@
</CENTER>
<P><BR></P>
<TABLE WIDTH="100%">
<TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/01-03i.jpg"><TD WIDTH="95%"><SMALL><I>Make sure you understand what really goes on when you insert a seemingly-innocuous function call into the time-critical portions of your code.</I></SMALL>
<TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>Make sure you understand what really goes on when you insert a seemingly-innocuous function call into the time-critical portions of your code.</I></SMALL>
</TABLE>
<P>In this case that means knowing how DOS and the C/C<SMALL>&#43;&#43;</SMALL> file-access libraries do their work. In other words, <I>know the territory</I>!</P>
<P><B>LISTING 1.4 L1-4.C</B></P>
@ -81,7 +81,7 @@ main(int argc, char *argv[]) {
<H4 ALIGN="LEFT"><A NAME="Heading10"></A><FONT COLOR="#000077">Know When It Matters</FONT></H4>
<P>The last section contained a particularly interesting phrase: <I>the time-critical portions of your code</I>. Time-critical portions of your code are those portions in which the speed of the code makes a significant difference in the overall performance of your program&#151;and by &#147;significant,&#148; I don&#146;t mean that it makes the code 100 percent faster, or 200 percent, or any particular amount at all, but rather that it makes the program more responsive and/or usable <I>from the user&#146;s perspective</I>.</P>
<P>Don&#146;t waste time optimizing non-time-critical code: set-up code, initialization code, and the like. Spend your time improving the performance of the code inside heavily-used loops and in the portions of your programs that directly affect response time. Notice, for example, that I haven&#146;t bothered to implement a version of the checksum program entirely in assembly; Listings 1.2 and 1.6 call assembly subroutines that handle the time-critical operations, but C is still used for checking command-line parameters, operning files, printing, and the like.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/01-04i.jpg"><TD WIDTH="95%"><SMALL><I>If you were to implement any of the listings in this chapter entirely in hand-optimized assembly, I suppose you might get a performance improvement of a few percent&#151;but I rather doubt you&#146;d get even that much, and you&#146;d sure as heck spend an awful lot of time for whatever meager improvement does result. Let C do what it does well, and use assembly only when it makes a perceptible difference.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>If you were to implement any of the listings in this chapter entirely in hand-optimized assembly, I suppose you might get a performance improvement of a few percent&#151;but I rather doubt you&#146;d get even that much, and you&#146;d sure as heck spend an awful lot of time for whatever meager improvement does result. Let C do what it does well, and use assembly only when it makes a perceptible difference.</I></SMALL>
</TABLE>
<P>Besides, we don&#146;t want to optimize until the design is refined to our satisfaction, and that won&#146;t be the case until we&#146;ve thought about other approaches.
</P>

View file

@ -37,7 +37,7 @@
</CENTER>
<P><BR></P>
<P>The third reason is often fallacious. C library functions are not always written in assembly, nor are they always particularly well-optimized. (In fact, they&#146;re often written for <I>portability</I>, which has nothing to do with optimization.) What&#146;s more, they&#146;re general-purpose functions, and often can be outperformed by well-but-not- brilliantly-written code that is well-matched to a specific task. As an example, consider Listing 1.5, which uses internal buffering to handle blocks of bytes at a time. Table 1.1 shows that Listing 1.5 is 2.5 to 4 times faster than Listing 1.4 (and as much as 49 times faster than Listing 1.1!), even though it uses no assembly at all.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/01-05i.jpg"><TD WIDTH="95%"><SMALL><I>Clearly, you can do well by using special-purpose C code in place of a C library function&#151;if you have a thorough understanding of how the C library function operates and exactly what your application needs done. Otherwise, you&#146;ll end up rewriting C library functions in C, which makes no sense at all.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>Clearly, you can do well by using special-purpose C code in place of a C library function&#151;if you have a thorough understanding of how the C library function operates and exactly what your application needs done. Otherwise, you&#146;ll end up rewriting C library functions in C, which makes no sense at all.</I></SMALL>
</TABLE>
<P><B>LISTING 1.5 L1-5.C</B></P>
<!-- CODE //-->

View file

@ -93,7 +93,7 @@ _ChecksumChunkendp
<P>Note that in Table 1.1, optimization makes little difference except in the case of Listing 1.5, where the design has been refined considerably. Execution time in the other cases is dominated by time spent in DOS and/or the C library, so optimization of the code you write is pretty much irrelevant. What&#146;s more, while the approximately two-times improvement we got by optimizing is not to be sneezed at, it pales against the up-to-50-times improvement we got by redesigning.
</P>
<P>By the way, the execution times even of Listings 1.6 and 1.7 are dominated by DOS disk access times. If a disk cache is enabled and the file to be checksummed is already in the cache, the assembly version is three times as fast as the C version. In other words, the inherent nature of this application limits the performance improvement that can be obtained via assembly. In applications that are more CPU-intensive and less disk-bound, particularly those applications in which string instructions and/or unrolled loops can be used effectively, assembly tends to be considerably faster relative to C than it is in this very specific case.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/01-06i.jpg"><TD WIDTH="95%"><SMALL><I>Don&#146;t get hung up on optimizing compilers or assembly language&#151;the best optimizer is between your ears.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>Don&#146;t get hung up on optimizing compilers or assembly language&#151;the best optimizer is between your ears.</I></SMALL>
</TABLE>
<P>All this is basically a way of saying: Know where you&#146;re going, know the territory, and know when it matters.
</P>

View file

@ -59,7 +59,7 @@ LoopTop:
<!-- END CODE SNIP //-->
<P>Now, it&#146;s hard to write code that&#146;s much faster than seven instructions, only one of which accesses memory, and most programmers would have called it a day at this point. Still, something bothered me, so I spent a bit of time going over the code again. Suddenly, the answer struck me&#151;the code was rotating each bit into place separately, so that a multibit rotation was being performed every time through the loop, for a total of four separate time-consuming multibit rotations!
</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/02-01i.jpg"><TD WIDTH="95%"><SMALL><I>While the instructions themselves were individually optimized, the overall approach did not make the best possible use of the instructions.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>While the instructions themselves were individually optimized, the overall approach did not make the best possible use of the instructions.</I></SMALL>
</TABLE>
<P>I changed the code to the following:
</P>

View file

@ -56,7 +56,7 @@
<P>In the PC world, you can never have enough knowledge, and every item you add to your store will make your programs better. Thorough familiarity with both the operating system APIs and BIOS interfaces is important; since those interfaces are well-documented and reasonably straightforward, my advice is to get a good book or two and bring yourself up to speed. Similarly, familiarity with the PC hardware is required. While that topic covers a lot of ground&#151;display adapters, keyboards, serial ports, printer ports, timer and DMA channels, memory organization, and more&#151;most of the hardware is well-documented, and articles about programming major hardware components appear frequently in the literature, so this sort of knowledge can be acquired readily enough.
</P>
<P>The single most critical aspect of the hardware, and the one about which it is hardest to learn, is the CPU. The x86 family CPUs have a complex, irregular instruction set, and, unlike most processors, they are neither straightforward nor wellregarding true code performance. What&#146;s more, assembly is so difficult to learn that most articles and books that present assembly code settle for code that just works, rather than code that pushes the CPU to its limits. In fact, since most articles and books are written for inexperienced assembly programmers, there is very little information of any sort available about how to generate high-quality assembly code for the x86 family CPUs. As a result, knowledge about programming them effectively is by far the hardest knowledge to gather. A good portion of this book is devoted to seeking out such knowledge.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/02-02i.jpg"><TD WIDTH="95%"><SMALL><I>Be forewarned, though: No matter how much you learn about programming the PC in assembly, there&#146;s always more to discover.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>Be forewarned, though: No matter how much you learn about programming the PC in assembly, there&#146;s always more to discover.</I></SMALL>
</TABLE>
<P><BR></P>
<CENTER>

View file

@ -41,7 +41,7 @@
</P>
<P>Basically, there are only two possible objectives to high-performance assembly programming: Given the requirements of the application, keep to a minimum either the number of processor cycles the program takes to run, or the number of bytes in the program, or some combination of both. We&#146;ll look at ways to achieve both objectives, but we&#146;ll more often be concerned with saving cycles than saving bytes, for the PC generally offers relatively more memory than it does processing horsepower. In fact, we&#146;ll find that two-to-three times performance improvements <I>over already tight assembly code</I> are often possible if we&#146;re willing to spend additional bytes in order to save cycles. It&#146;s not always desirable to use such techniques to speed up code, due to the heavy memory requirements&#151;but it is almost always <I>possible</I>.</P>
<P>You will notice that my short list of objectives for high-performance assembly programming does not include traditional objectives such as easy maintenance and speed of development. Those are indeed important considerations&#151;to persons and companies that develop and distribute software. People who actually <I>buy</I> software, on the other hand, care only about how well that software performs, not how it was developed nor how it is maintained. These days, developers spend so much time focusing on such admittedly important issues as code maintainability and reusability, source code control, choice of development environment, and the like that they often forget rule #1: From the user&#146;s perspective, <I>performance is fundamental</I>.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/02-03i.jpg"><TD WIDTH="95%"><SMALL><I>Comment your code, design it carefully, and write non-time-critical portions in a high-level language, if you wish&#151;but when you write the portions that interact with the user and/or affect response time, performance must be your paramount objective, and assembly is the path to that goal</I>.</SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>Comment your code, design it carefully, and write non-time-critical portions in a high-level language, if you wish&#151;but when you write the portions that interact with the user and/or affect response time, performance must be your paramount objective, and assembly is the path to that goal</I>.</SMALL>
</TABLE>
<P>Knowledge of the sort described earlier is absolutely essential to fulfilling either of the objectives of assembly programming. What that knowledge doesn&#146;t do by itself is meet the need to write code that both performs to the requirements of the application at hand and also operates as efficiently as possible in the PC environment. Knowledge makes that possible, but your programming instincts make it happen. And it is that intuitive, on-the-fly integration of a program specification and a sea of facts about the PC that is the heart of the Zen-class assembly optimization.
</P>

View file

@ -46,7 +46,7 @@
<H3><A NAME="Heading3"></A><FONT COLOR="#000077">The Costs of Ignorance</FONT></H3>
<P>As diligent as the author had been, he had nonetheless committed a cardinal sin of x86 assembly language programming: He had assumed that the information available to him was both correct and complete. While the execution times provided by Intel for its processors are indeed correct, they are incomplete; the other&#151;and often more important&#151;part of code performance is instruction <I>fetch</I> time, a topic to which I will return in later chapters.</P>
<P>Had the author taken the time to measure the true performance of his code, he wouldn&#146;t have put his reputation on the line with relatively low-performance code. What&#146;s more, had he actually measured the performance of his code and found it to be unexpectedly slow, curiosity might well have led him to experiment further and thereby add to his store of reliable information about the CPU.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/03-01i.jpg"><TD WIDTH="95%" VALIGN="TOP"><I><SMALL>There you have an important tenet of assembly language optimization: After crafting the best code possible, check it in action to see if it&#146;s really doing what you think it is. If it&#146;s not behaving as expected, that&#146;s all to the good, since solving mysteries is the path to knowledge. You&#146;ll learn more in this way, I assure you, than from any manual or book on assembly language.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%" VALIGN="TOP"><I><SMALL>There you have an important tenet of assembly language optimization: After crafting the best code possible, check it in action to see if it&#146;s really doing what you think it is. If it&#146;s not behaving as expected, that&#146;s all to the good, since solving mysteries is the path to knowledge. You&#146;ll learn more in this way, I assure you, than from any manual or book on assembly language.</I></SMALL>
</TABLE>
<P><I>Assume nothing</I>. I cannot emphasize this strongly enough&#151;when you care about performance, do your best to improve the code and then <I>measure</I> the improvement. If you don&#146;t measure performance, you&#146;re just guessing, and if you&#146;re guessing, you&#146;re not very likely to write top-notch code.</P>
<P>Ignorance about true performance can be costly. When I wrote video games for a living, I spent days at a time trying to wring more performance from my graphics drivers. I rewrote whole sections of code just to save a few cycles, juggled registers, and relied heavily on blurry-fast register-to-register shifts and adds. As I was writing my last game, I discovered that the program ran perceptibly faster if I used look-up tables instead of shifts and adds for my calculations. It <I>shouldn&#146;t</I> have run faster, according to my cycle counting, but it did. In truth, instruction fetching was rearing its head again, as it often does, and the fetching of the shifts and adds was taking as much as four times the nominal execution time of those instructions.</P>

View file

@ -39,7 +39,7 @@
<H4 ALIGN="LEFT"><A NAME="Heading10"></A><FONT COLOR="#000077">Official Execution Times Are Only Part of the Story</FONT></H4>
<P>The sequence of 5 <B>SHR</B> instructions in the last example is 10 bytes long. That means that it can never execute in less than 24 cycles even if the 4-byte prefetch queue is full when it starts, since 6 instruction bytes would still remain to be fetched, at 4 cycles per fetch. If the prefetch queue is empty at the start, the sequence <I>could</I> take 40 cycles. In short, thanks to instruction fetching, the code won&#146;t run at its documented speed, and could take up to four times longer than it is supposed to.</P>
<P>Why does Intel document Execution Unit execution time rather than overall instruction execution time, which includes both instruction fetch time and Execution Unit (EU) execution time? Well, instruction fetching isn&#146;t performed as part of instruction execution by the Execution Unit, but instead is carried on in parallel by the Bus Interface Unit (BIU) whenever the external data bus isn&#146;t in use or whenever the EU runs out of instruction bytes to execute. Sometimes the BIU is able to use spare bus cycles to prefetch instruction bytes before the EU needs them, so in those cases instruction fetching takes no time at all, practically speaking. At other times the EU executes instructions faster than the BIU can fetch them, and instruction fetching then becomes a significant part of overall execution time. As a result, <I>the effective fetch time for a given instruction varies greatly depending on the code mix preceding that instruction.</I> Similarly, the state in which a given instruction leaves the prefetch queue affects the overall execution time of the following instructions.</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/04-01i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>In other words, while the execution time for a given instruction is constant, the fetch time for that instruction depends heavily on the context in which the instruction is executing&#151;the amount of prefetching the preceding instructions allowed&#151;and can vary from a full 4 cycles per instruction byte to no time at all.</I></SMALL>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>In other words, while the execution time for a given instruction is constant, the fetch time for that instruction depends heavily on the context in which the instruction is executing&#151;the amount of prefetching the preceding instructions allowed&#151;and can vary from a full 4 cycles per instruction byte to no time at all.</I></SMALL>
</TABLE>
<P>As we&#146;ll see later, other cycle-eaters, such as DRAM refresh and display memory wait states, can cause prefetching variations even during different executions of the same code sequence. Given that, it&#146;s meaningless to talk about the prefetch time of a given instruction except in the context of a specific code sequence.
</P>

View file

@ -74,7 +74,7 @@
<P>As you can see, it&#146;s easy to be drawn into thinking you&#146;re saving cycles when you&#146;re not. You can only improve the performance of a specific bit of code by reducing the factor&#151;either instruction fetch time or execution time, or sometimes a mix of the two&#151;that&#146;s limiting the performance of that code.
</P>
<P>In case you missed it in all the excitement, the variability of prefetching means that our method of testing performance by executing 1,000 instructions in a row by no means produces &#147;true&#148; instruction execution times, any more than the official execution times in the Intel manuals are &#147;true&#148; times. The fact of the matter is that a given instruction takes <I>at least</I> as long to execute as the time given for it in the Intel manuals, but may take as much as 4 cycles per byte longer, depending on the state of the prefetch queue when the preceding instruction ends.</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/04-02i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>The only true execution time for an instruction is a time measured in a certain context, and that time is meaningful only in that context.</I></SMALL>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>The only true execution time for an instruction is a time measured in a certain context, and that time is meaningful only in that context.</I></SMALL>
</TABLE>
<P>What we <I>really</I> want is to know how long useful working code takes to run, not how long a single instruction takes, and the Zen timer gives us the tool we need to gather that information. Granted, it would be easier if we could just add up neatly documented instruction execution times&#151;but that&#146;s not going to happen. Without actually measuring the performance of a given code sequence, you simply don&#146;t know how fast it is. For crying out loud, even the people who <I>designed</I> the 8088 at Intel couldn&#146;t tell you exactly how quickly a given 8088 code sequence executes on the PC just by looking at it! Get used to the idea that execution times are only meaningful in context, learn the rules of thumb in this book, and use the Zen timer to measure your code.</P>
<H4 ALIGN="LEFT"><A NAME="Heading12"></A><FONT COLOR="#000077">Approximating Overall Execution Times</FONT></H4>

View file

@ -37,7 +37,7 @@
</CENTER>
<P><BR></P>
<TABLE WIDTH="100%">
<TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/04-03i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>A line-drawing subroutine, which executes perhaps a dozen instructions for each display memory access, generally loses less performance to the display adapter cycle-eater than does a block-copy or scrolling subroutine that uses <B>REP MOVS</B> instructions. Scaled and three-dimensional graphics, which spend a great deal of time performing calculations (often using very slow floating-point arithmetic), tend to suffer less.</I></SMALL>
<TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>A line-drawing subroutine, which executes perhaps a dozen instructions for each display memory access, generally loses less performance to the display adapter cycle-eater than does a block-copy or scrolling subroutine that uses <B>REP MOVS</B> instructions. Scaled and three-dimensional graphics, which spend a great deal of time performing calculations (often using very slow floating-point arithmetic), tend to suffer less.</I></SMALL>
</TABLE>
<P>In addition, code that accesses display memory infrequently tends to suffer only about half of the maximum display memory wait states, because on average such code will access display memory halfway between one available display memory access slot and the next. As a result, code that accesses display memory less intensively than the code in Listing 4.11 will on average lose 4 or 5 rather than 8-plus cycles to the display adapter cycle-eater on each memory access.
</P>
@ -45,7 +45,7 @@
<H4 ALIGN="LEFT"><A NAME="Heading22"></A><FONT COLOR="#000077">What to Do about the Display Adapter Cycle-Eater?</FONT></H4>
<P>What can we do about the display adapter cycle-eater? Well, we can minimize display memory accesses whenever possible. In particular, we can try to avoid read/modify/write display memory operations of the sort used to mask individual pixels and clip images. Why? Because read/modify/write operations require two display memory accesses (one read and one write) each time display memory is manipulated. Instead, we should try to use writes of the sort that set all the pixels in a given byte of display memory at once, since such writes don&#146;t require accompanying read accesses. The key here is that only half as many display memory accesses are required to write a byte to display memory as are required to read a byte from display memory, mask part of it off and alter the rest, and write the byte back to display memory. Half as many display memory accesses means half as many display memory wait states.
</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/04-04i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Moreover, 486s and Pentiums, as well as recent Super VGAs, employ write-caching schemes that make display memory writes considerably faster than display memory reads.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Moreover, 486s and Pentiums, as well as recent Super VGAs, employ write-caching schemes that make display memory writes considerably faster than display memory reads.</I></SMALL>
</TABLE>
<P>Along the same line, the display adapter cycle-eater makes the popular exclusive-OR animation technique, which requires paired reads and writes of display memory, less-than-ideal for the PC. Exclusive-OR animation should be avoided in favor of simply writing images to display memory whenever possible.
</P>

View file

@ -46,7 +46,7 @@
<P>In Chapter 1, I briefly discussed using <I>restartable blocks</I>. This, you might remember, is the process of handling in chunks data sets too large to fit in memory so that they can be processed just about as fast as if they did fit in memory. The restartable block approach is very fast but is relatively difficult to program.</P>
<P>At the opposite end of the spectrum lies byte-by-byte processing, whereby DOS (or, in less extreme cases, a group of library functions) is allowed to do all the hard work, so that you only have to deal with one byte at a time. Byte-by-byte processing is easy to program but can be extremely slow, due to the vast overhead that results from invoking DOS each time a byte must be processed.</P>
<P>Sound familiar? It should. I moved via the byte-by-byte approach, and the overhead of driving back and forth made for miserable performance. Renting a truck (the restartable block approach) would have required more effort and forethought, but would have paid off handsomely.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/05-01i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>The easy, familiar approach often has nothing in its favor except that it requires less thinking; not a great virtue when writing high-performance code&#151;or when moving.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>The easy, familiar approach often has nothing in its favor except that it requires less thinking; not a great virtue when writing high-performance code&#151;or when moving.</I></SMALL>
</TABLE>
<P>And with that, let&#146;s look at a fairly complex application of restartable blocks.
</P>

View file

@ -37,7 +37,7 @@
</CENTER>
<P><BR></P>
<P>We could put a zero byte at the end of our buffer to allow <B>strstr()</B> to work, but why bother? The <B>strstr()</B> function must spend time either checking for the end of the string being searched or determining the length of that string&#151;wasted effort given that we already know exactly how long our search buffer is. Even if a given <B>strstr()</B> implementation is well-written, its performance will suffer, at least for our application, from unnecessary overhead.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/05-02i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>This illustrates why you shouldn&#146;t think of C/C<SMALL>&#43;&#43;</SMALL> library functions as black boxes; understand what they do and try to figure out how they do it, and relate that to their performance in the context you&#146;re interested in.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>This illustrates why you shouldn&#146;t think of C/C<SMALL>&#43;&#43;</SMALL> library functions as black boxes; understand what they do and try to figure out how they do it, and relate that to their performance in the context you&#146;re interested in.</I></SMALL>
</TABLE>
<H3><A NAME="Heading5"></A><FONT COLOR="#000077">Brute-Force Techniques</FONT></H3>
<P>Given that no C/C<SMALL>&#43;&#43;</SMALL> library function meets our needs precisely, an obvious alternative approach is the brute-force technique that uses <B>memcmp()</B> to compare <I>every</I> potential matching location in the buffer to the string we&#146;re searching for, as illustrated in Figure 5.1.</P>

View file

@ -39,7 +39,7 @@
<P>It might also be worth converting the search engine to assembly for searches performed entirely in memory; with the overhead of file access eliminated, improvements in search-engine performance would translate directly into significantly faster overall performance. One such application that would have much the same structure as Listing 5.1 would be searching through expanded memory buffers, and another would be searching through huge (segment-spanning) buffers.
</P>
<P>And so we find, as we so often will, that optimization is definitely not a cut-and-dried matter, and that there is no such thing as a single &#147;best&#148; approach.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/05-03i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>You must know what your application will typically do, and you must know whether you&#146;re more concerned with average or worst-case performance before you can decide how best to speed up your program&#151;and, indeed, whether speeding it up is worth doing at all.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>You must know what your application will typically do, and you must know whether you&#146;re more concerned with average or worst-case performance before you can decide how best to speed up your program&#151;and, indeed, whether speeding it up is worth doing at all.</I></SMALL>
</TABLE>
<P>By the way, don&#146;t think that just because very large block sizes don&#146;t much improve performance, it wasn&#146;t worth using restartable blocks in Listing 5.1. Listing 5.1 runs more than three times more slowly with a block size of 32 bytes than with a block size of 4K, and any byte-by-byte approach would surely be slower still, due to the overhead of repeated calls to DOS and/or the C stream I/O library.
</P>
@ -48,7 +48,7 @@
<P>I&#146;ve explained two important lessons: Know when it&#146;s worth optimizing further, and use restartable blocks to process large data sets as a series of blocks, with each block handled at high speed. The first lesson is less obvious than it seems.
</P>
<P>When I set out to write this chapter, I fully intended to write an assembly language version of Listing 5.1, and I expected the assembly version to be much faster. When I actually looked at where execution time was going (which I did by modifying the program to remove the calls to the <B>read()</B> function, but a code profiler could be used to do the same thing much more easily), I found that the best code in the world wouldn&#146;t make much difference.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/05-04i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>When you try to speed up code, take a moment to identify the hot spots in your program so that you know where optimization is needed and whether it will make a significant difference before you invest your time.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>When you try to speed up code, take a moment to identify the hot spots in your program so that you know where optimization is needed and whether it will make a significant difference before you invest your time.</I></SMALL>
</TABLE>
<P>As for restartable blocks: Here we tackled a considerably more complex application of restartable blocks than we did in Chapter 1&#151;which turned out not to be so difficult after all. Don&#146;t let irregularities in the programming tasks you tackle, such as strings that span blocks, fluster you into settling for easy, general&#151;and slow&#151;solutions. Focus on making the inner loop&#151;the code that handles each block&#151;as efficient as possible, then structure the rest of your code to support the inner loop.
</P>

View file

@ -132,7 +132,7 @@ add ebx,edx
<P>and would in either case require the destruction of the contents of another register.
</P>
<P>Multiplying a 32-bit value by a non-power-of-two multiplier in just 2 cycles is a pretty neat trick, even though it works only on a 386 or 486.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/06-01i.jpg"><TD WIDTH="95%"><SMALL><I>The full list of values that <B>LEA</B> can multiply a register by on a 386 or 486 is: 2, 3, 4, 5, 8, and 9. That list doesn&#146;t include every multiplier you might want, but it covers some commonly used ones, and the performance is hard to beat.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>The full list of values that <B>LEA</B> can multiply a register by on a 386 or 486 is: 2, 3, 4, 5, 8, and 9. That list doesn&#146;t include every multiplier you might want, but it covers some commonly used ones, and the performance is hard to beat.</I></SMALL>
</TABLE>
<P>I&#146;d like to extend my thanks to Duane Strong of Metagraphics for his help in brainstorming uses for the 386 version of <B>LEA</B> and for pointing out the complications of 486 instruction timings.</P><P><BR></P>
<CENTER>

View file

@ -50,16 +50,16 @@
<P>68 degrees is warm for an uninsulated third floor in Buffalo in the dead of winter. Damn warm. It is not, however, particularly warm for a sauna. Eventually someone acknowledges the obvious and allows that it might have been a stupid idea after all, and everyone agrees, and they shut off the heater and leave, each no doubt offering silent thanks that they had gotten out of this without any incidents requiring major surgery.</P>
<P>And so we see that the best idea in the world can fail for lack of either proper design or adequate horsepower. The primary cause of the Great Buffalo Sauna Fiasco was a lack of horsepower; the gas heater was flat-out undersized. This is analogous to trying to write programs that incorporate features like bitmapped text and searching of multisegment buffers without using high-performance assembly language. Any PC language can perform just about any function you can think of&#151;eventually. That heater would eventually have heated the room to 110 degrees, too&#151;along about the first of June or so.</P>
<P>The Great Buffalo Sauna Fiasco also suffered from fundamental design flaws. A more powerful heater would indeed have made the room hotter&#151;and might well have burned the house down in the process. Likewise, proper algorithm selection and good design are fundamental to performance. The extra horsepower a superb assembly language implementation gives a program is worth bothering with only in the context of a good design.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/07-01i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Assembly language optimization is a small but crucial corner of the PC programming world. Use it sparingly and only within the framework of a good design&#151;but ignore it and you may find various portions of your anatomy out in the cold.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Assembly language optimization is a small but crucial corner of the PC programming world. Use it sparingly and only within the framework of a good design&#151;but ignore it and you may find various portions of your anatomy out in the cold.</I></SMALL>
</TABLE>
<P>So, drawing fortitude from the knowledge that our quest is a pure and worthy one, let&#146;s resume our exploration of assembly language instructions with hidden talents and instructions with well-known talents that are less than they appear to be. In the process, we&#146;ll come to see that there is another, very important optimization level between the algorithm/design level and the cycle-counting/individual instruction level. I&#146;ll call this middle level <I>local optimization;</I> it involves focusing on optimizing sequences of instructions rather than individual instructions, all with an eye to implementing designs as efficiently as possible given the capabilities of the x86 family instruction set.</P>
<P>And yes, in case you&#146;re wondering, the above story is indeed true. Was I there? Let me put it this way: If I were, I&#146;d never admit it!</P>
<H4 ALIGN="LEFT"><A NAME="Heading3"></A><FONT COLOR="#000077">When LOOP Is a Bad Idea</FONT></H4>
<P>Let&#146;s examine first an instruction that is less than it appears to be: <B>LOOP</B>. There&#146;s no mystery about what <B>LOOP</B> does; it decrements CX and branches if CX doesn&#146;t decrement to zero. It&#146;s so beautifully suited to the task of counting down loops that any experienced x86 programmer instinctively stuffs the loop count in CX and reaches for <B>LOOP</B> when setting up a loop. That&#146;s fine&#151;<B>LOOP</B> does, of course, work as advertised&#151;but there is one problem:</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/07-02i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>On half of the processors in the x86 family, <B>LOOP</B> is slower than <B>DEC CX</B> followed by <B>JNZ</B>. (Granted, <B>DEC CX/JNZ</B> isn&#146;t precisely equivalent to <B>LOOP,</B> because <B>DEC</B> alters the flags and LOOP doesn&#146;t, but in most situations they&#146;re comparable.)</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>On half of the processors in the x86 family, <B>LOOP</B> is slower than <B>DEC CX</B> followed by <B>JNZ</B>. (Granted, <B>DEC CX/JNZ</B> isn&#146;t precisely equivalent to <B>LOOP,</B> because <B>DEC</B> alters the flags and LOOP doesn&#146;t, but in most situations they&#146;re comparable.)</I></SMALL>
</TABLE>
<P>How can this be? Don&#146;t ask me, ask Intel. On the 8088 and 80286, <B>LOOP</B> is indeed faster than <B>DEC CX/JNZ</B> by a cycle, and <B>LOOP</B> is generally a little faster still because it&#146;s a byte shorter and so can be fetched faster. On the 386, however, things change; <B>LOOP</B> is two cycles <I>slower</I> than <B>DEC/JNZ,</B> and the fetch time for one extra byte on even an uncached 386 generally isn&#146;t significant. (Remember that the 386 fetches four instruction bytes at a pop.) <B>LOOP</B> is three cycles slower than <B>DEC/JNZ</B> on the 486, and the 486 executes instructions in so few cycles that those three cycles mean that <B>DEC/JNZ</B> is nearly <I>twice</I> as fast as <B>LOOP</B>. Then, too, unlike <B>LOOP, DEC</B> doesn&#146;t require that <B>CX</B> be used, so the <B>DEC/JNZ</B> solution is both faster and more flexible on the 386 and 486, and on the Pentium as well. (By the way, all this is not just theory; I&#146;ve timed the relative performances of <B>LOOP</B> and <B>DEC CX/JNZ</B> on a cached 386, and LOOP really is slower.)</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/07-03i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Things are stranger still for <B>LOOP</B>&#146;s relative <B>JCXZ,</B> which branches if and only if CX is zero. <B>JCXZ</B> is faster than <B>AND CX,CX/JZ</B> on the 8088 and 80286, and equivalent on the 80386&#151;but is about twice as slow on the 486!</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Things are stranger still for <B>LOOP</B>&#146;s relative <B>JCXZ,</B> which branches if and only if CX is zero. <B>JCXZ</B> is faster than <B>AND CX,CX/JZ</B> on the 8088 and 80286, and equivalent on the 80386&#151;but is about twice as slow on the 486!</I></SMALL>
</TABLE>
<P><BR></P>
<CENTER>

View file

@ -60,7 +60,7 @@ jz SkipLoop ;If field is 0, don&#146;t bother
<P>One level at which assembly language programming pays off handsomely is that of <I>local optimization;</I> that is, selecting the best <I>sequence</I> of instructions for a task. The key to local optimization is viewing the 80x86 instruction set as a set of building blocks, each with unique characteristics. Your job is to sequence those blocks so that they perform well. It doesn&#146;t matter what the instructions are intended to do or what their names are; all that matters is what they <I>do.</I></P>
<P>Our discussion of <B>LOOP</B> versus <B>DEC/JNZ</B> is an excellent example of optimization by cycle counting. It&#146;s worth knowing, but once you&#146;ve learned it, you just routinely use <B>DEC/JNZ</B> at the bottom of loops in 386/486-specific code, and that&#146;s that. Besides, you&#146;ll save at most a few cycles each time, and while that helps a little, it&#146;s not going to make all <I>that</I> much difference.</P>
<P>Now let&#146;s step back for a moment, and with no preconceptions consider what the x86 instruction set can do for us. The bulk of the time with both <B>LOOP</B> and <B>DEC/JNZ</B> is taken up by branching, which just happens to be one of the slowest aspects of every processor in the x86 family, and the rest is taken up by decrementing the count register and checking whether it&#146;s zero. There may be ways to perform those tasks a little faster by selecting different instructions, but they can get only so fast, and branching can&#146;t even get all that fast.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/07-04i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>The trick, then, is not to find the fastest way to decrement a count and branch conditionally, but rather to figure out how to accomplish the same result without decrementing or branching as often. Remember the Kobiyashi Maru problem in</I> Star Trek<I>?The same principle applies here: Redefine the problem to one that offers better solutions.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>The trick, then, is not to find the fastest way to decrement a count and branch conditionally, but rather to figure out how to accomplish the same result without decrementing or branching as often. Remember the Kobiyashi Maru problem in</I> Star Trek<I>?The same principle applies here: Redefine the problem to one that offers better solutions.</I></SMALL>
</TABLE>
<P>Consider Listing 7.1, which searches a buffer until either the specified byte is found, a zero byte is found, or the specified number of characters have been checked. Such a function would be useful for scanning up to a maximum number of characters in a zero-terminated buffer. Listing 7.1, which uses <B>LOOP</B> in the main loop, performs a search of the sample string for a period (&#145;.&#146;) in 170 &#181;s on a 20 MHz cached 386.</P>
<P>When the <B>LOOP</B> in Listing 7.1 is replaced with <B>DEC CX/JNZ,</B> performance improves to 168 &#181;s, less than 2 percent faster than Listing 7.1. Actually, instruction fetching, instruction alignment, cache characteristics, or something similar is affecting these results; I&#146;d expect a slightly larger improvement&#151;around 7 percent&#151;but that&#146;s the most that counting cycles could buy us in this case. (All right, already; <B>LOOPNZ</B> could be used at the bottom of the loop, and other optimizations are surely possible, but all that won&#146;t add up to anywhere near the benefits we&#146;re about to see from local optimization, and that&#146;s the whole point.)</P><P><BR></P>

View file

@ -170,7 +170,7 @@ SearchMaxLengthendp
</PRE>
<!-- END CODE //-->
<P>How much difference? Listing 7.2 runs in 121 &#181;s&#151;40 percent faster than Listing 7.1, even though Listing 7.2 still uses <B>LOOP</B> rather than <B>DEC CX/JNZ.</B> (The loop in Listing 7.2 could be unrolled further, too; it&#146;s just a question of how much more memory you want to trade for ever-decreasing performance benefits.) That&#146;s typical of local optimization; it won&#146;t often yield the order-of-magnitude improvements that algorithmic improvements can produce, but it can get you a critical 50 percent or 100 percent improvement when you&#146;ve exhausted all other avenues.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/07-05i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>The point is simply this: You can gain far more by stepping back a bit and thinking of the fastest overall way for the CPU to perform a task than you can by saving a cycle here or there using different instructions. Try to think at the level of sequences of instructions rather than individual instructions, and learn to treat x86 instructions as building blocks with unique characteristics rather than as instructions dedicated to specific tasks.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>The point is simply this: You can gain far more by stepping back a bit and thinking of the fastest overall way for the CPU to perform a task than you can by saving a cycle here or there using different instructions. Try to think at the level of sequences of instructions rather than individual instructions, and learn to treat x86 instructions as building blocks with unique characteristics rather than as instructions dedicated to specific tasks.</I></SMALL>
</TABLE>
<P><BR></P>
<CENTER>

View file

@ -79,11 +79,11 @@ BIT_PATTERN=BIT_PATTERN SHL 1
</PRE>
<!-- END CODE SNIP //-->
<TABLE WIDTH="100%">
<TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/07-06i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Besides illustrating the advantages of local optimization, this example also shows that it generally pays to precalculate results; this is often done at or before assembly time, but precalculated tables can also be built at run time. This is merely one aspect of a fundamental optimization rule: Move as much work as possible out of your critical code by whatever means necessary.</I></SMALL>
<TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Besides illustrating the advantages of local optimization, this example also shows that it generally pays to precalculate results; this is often done at or before assembly time, but precalculated tables can also be built at run time. This is merely one aspect of a fundamental optimization rule: Move as much work as possible out of your critical code by whatever means necessary.</I></SMALL>
</TABLE>
<H4 ALIGN="LEFT"><A NAME="Heading9"></A><FONT COLOR="#000077">NOT Flips Bits&#151;Not Flags</FONT></H4>
<P>The <B>NOT</B> instruction flips all the bits in the operand, from 0 to 1 or from 1 to 0. That&#146;s as simple as could be, but <B>NOT</B> nonetheless has a minor but interesting talent: It doesn&#146;t affect the flags. That can be irritating; I once spent a good hour tracking down a bug caused by my unconscious assumption that <B>NOT</B> does set the flags. After all, every other arithmetic and logical instruction sets the flags; why not <B>NOT</B>? Probably because <B>NOT</B> isn&#146;t considered to be an arithmetic or logical instruction at all; rather, it&#146;s a data manipulation instruction, like <B>MOV</B> and the various rotates. (These are <B>RCR, RCL, ROR,</B> and <B>ROL,</B> which affect only the Carry and Overflow flags.) NOT is often used for tasks, such as flipping masks, where there&#146;s no reason to test the state of the result, and in that context it can be handy to keep the flags unmodified for later testing.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/07-07i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Besides, if you want to <B>NOT</B> an operand and set the flags in the process, you can just <B>XOR</B> it with -1. Put another way, the only functional difference between <B>NOT AX</B> and <B>XOR AX,0FFFFH</B> is that <B>XOR</B> modifies the flags and <B>NOT</B> doesn&#146;t.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Besides, if you want to <B>NOT</B> an operand and set the flags in the process, you can just <B>XOR</B> it with -1. Put another way, the only functional difference between <B>NOT AX</B> and <B>XOR AX,0FFFFH</B> is that <B>XOR</B> modifies the flags and <B>NOT</B> doesn&#146;t.</I></SMALL>
</TABLE>
<P>The x86 instruction set offers many ways to accomplish almost any task. Understanding the subtle distinctions between the instructions&#151;whether and which flags are set, for example&#151;can be critical when you&#146;re trying to optimize a code sequence and you&#146;re running out of registers, or when you&#146;re trying to minimize branching.
</P>
@ -121,7 +121,7 @@ LOOP_TOP:
<!-- END CODE //-->
<P>It&#146;s not that the Listing 7.6 approach is necessarily better or worse; that depends on the processor and the situation. The Listing 7.6 approach is <I>different,</I> and if you understand the differences, you&#146;ll be able to choose the best approach for whatever code you happen to write. (<B>DEC</B> has the same property of preserving the Carry flag, by the way.)</P>
<P>There are a couple of interesting aspects to the last example. First, note that <B>LOOP</B> doesn&#146;t affect any flags at all; this allows the Carry flag to remain unchanged from one addition to the next. Not altering the arithmetic flags is a common characteristic of program control instructions (as opposed to arithmetic and logical instructions like <B>SUB</B> and <B>AND,</B> which do alter the flags).</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/07-08i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>The rule is not that the arithmetic flags change whenever the CPU performs a calculation; rather, the flags change whenever you execute an arithmetic, logical, or flag control (such as <B>CLC</B> to clear the Carry flag) instruction.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>The rule is not that the arithmetic flags change whenever the CPU performs a calculation; rather, the flags change whenever you execute an arithmetic, logical, or flag control (such as <B>CLC</B> to clear the Carry flag) instruction.</I></SMALL>
</TABLE>
<P>Not only do <B>LOOP</B> and <B>JCXZ</B> not alter the flags, but <B>REP MOVS</B>, which counts down CX to 0, doesn&#146;t affect the flags either.</P>
<P>The other interesting point about the last example is the use of <B>LAHF</B> and <B>SAHF,</B> which transfer the low byte of the FLAGS register to and from AH, respectively. These instructions were created to help provide compatibility with the 8080&#146;s (that&#146;s <I>8080</I>, not <I>8088</I>) <B>PUSH</B> <B>PSW</B> and <B>POP PSW</B> instructions, but turn out to be compact (one byte) instructions for saving and restoring the arithmetic flags. A word of caution, however: <B>SAHF</B> restores the Carry, Zero, Sign, Auxiliary Carry, and Parity flags&#151;but <I>not</I> the Overflow flag, which resides in the high byte of the FLAGS register. Also, be aware that <B>LAHF</B> and <B>SAHF</B> provide a fast way to preserve the flags on an 8088 but are relatively slow instructions on the 486 and Pentium.</P>

View file

@ -44,7 +44,7 @@
<P>&#147;Well,&#148; said Jeff, &#147;I think it suffered in the translation from the French.&#148;</P>
<P>Ah-ha! Mystery solved. Apparently everyone but me knew that it was translated from French, and that novelty undoubtedly made the song a big hit. The translation was also surely responsible for the sappy lyrics; dollars to donuts that the original French lyrics were stronger.</P>
<P>Which brings us without missing a beat to this chapter&#146;s theme, speeding up C with assembly language. When you seek to speed up a C program by converting selected parts of it (generally no more than a few functions) to assembly language, make sure you end up with high-performance assembly language code, not fine-tuned C code. Compilers like Microsoft C/C<SMALL>&#43;&#43;</SMALL> and Watcom C are by now pretty good at fine-tuning C code, and you&#146;re not likely to do much better by taking the compiler&#146;s assembly language output and tweaking it.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/08-01i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>To make the process of translating C code to assembly language worth the trouble, you must ignore what the compiler does and design your assembly language code from a pure assembly language perspective. With a merely adequate translation, you risk laboring mightily for little or no reward.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>To make the process of translating C code to assembly language worth the trouble, you must ignore what the compiler does and design your assembly language code from a pure assembly language perspective. With a merely adequate translation, you risk laboring mightily for little or no reward.</I></SMALL>
</TABLE>
<P>Apropos of which, when was the last time you heard of Terry Jacks?
</P>
@ -53,7 +53,7 @@
<P>Before I discuss Rule 1 further, let me mention rule number 0: <I>Only optimize where it matters.</I> The bulk of execution time in any program is spent in a very small portion of the code, and most code beyond that small portion doesn&#146;t have any perceptible impact on performance. Unless you&#146;re supremely concerned with code size (an area in which assembly-only programs can excel), I&#146;d suggest that you write most of your code in C and reserve assembly for the truly critical sections of your code; that&#146;s the formula that I find gives the most bang for the buck.</P>
<P>This is not to say that complete programs shouldn&#146;t be <I>designed</I> with optimized assembly language in mind. As you&#146;ll see shortly, orienting your data structures towards assembly language can be a salubrious endeavor indeed, even if most of your code is in C. When it comes to actually optimizing code and/or converting it to assembly, though, do it only where it matters. Get a profiler&#151;and use it!</P>
<P>Also make it a point to concentrate on refining your program design and algorithmic approach at the conceptual and/or C levels before doing any assembly language optimization.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/08-02i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Assembly language optimization is the final and far from the only step in the optimization chain, and as such should be performed last; converting to assembly too soon can lock in your code before the design is optimal. At the very least, conversion to assembly tends to make future changes and debugging more difficult, slowing you down and limiting your options.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Assembly language optimization is the final and far from the only step in the optimization chain, and as such should be performed last; converting to assembly too soon can lock in your code before the design is optimal. At the very least, conversion to assembly tends to make future changes and debugging more difficult, slowing you down and limiting your options.</I></SMALL>
</TABLE>
<H3><A NAME="Heading4"></A><FONT COLOR="#000077">Don&#146;t Call Your Functions on Me, Baby</FONT></H3>
<P>In order to think differently from a compiler, you must understand both what compilers and C programmers tend to do and how that differs from what assembly language does well. In this pursuit, it can be useful to examine the code your compiler generates, either by viewing the code in a debugger or by having the compiler generate an assembly language output file. (The latter is done with /Fa or /Fc in Microsoft C/C<SMALL>&#43;&#43;</SMALL> and -S in Borland C<SMALL>&#43;&#43;</SMALL>.)</P>

View file

@ -43,14 +43,14 @@
<H3><A NAME="Heading6"></A><FONT COLOR="#000077">Torn Between Two Segments</FONT></H3>
<P>C compilers are not terrific at handling segments. Some compilers can efficiently handle a single far pointer used in a loop by leaving ES set for the duration of the loop. But two far pointers used in the same loop confuse every compiler I&#146;ve seen, causing the full segment:offset address to be reloaded each time either pointer is used.
</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/08-03i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>This particularly affects performance in 286 protected mode (under OS/2 1.X or the Rational DOS Extender, for example) because segment loads in protected mode take a minimum of 17 cycles, versus a mere 2 cycles in real mode. </I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>This particularly affects performance in 286 protected mode (under OS/2 1.X or the Rational DOS Extender, for example) because segment loads in protected mode take a minimum of 17 cycles, versus a mere 2 cycles in real mode. </I></SMALL>
</TABLE>
<P>In assembly language you have full control over segments. Use it, and, if necessary, reorganize your code to minimize segment loading.
</P>
<H4 ALIGN="LEFT"><A NAME="Heading7"></A><FONT COLOR="#000077">Why Speeding Up Is Hard to Do</FONT></H4>
<P>You might think that the most obvious advantage assembly language has over C is that it allows the use of all forms of instructions and all registers in all ways, whereas C compilers tend to use a subset of registers and instructions in a limited number of ways. Yes and no. It&#146;s true that C compilers typically don&#146;t generate instructions such as <B>XLAT,</B> rotates, or the string instructions. On the other hand, <B>XLAT</B> and rotates are useful in a limited set of circumstances, and string instructions <I>are</I> used in the C library functions. In fact, C library code is likely to be carefully optimized by experts, and may be much better than equivalent code you&#146;d produce yourself.</P>
<P>Am I saying that C compilers produce better code than you do? No, I&#146;m saying that they <I>can,</I> unless you use assembly language properly. Writing code in assembly language rather than C guarantees nothing.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/08-04i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>You can write good assembly, bad assembly, or assembly that is virtually indistinguishable from compiled code; you are more likely than not to write the latter if you think that optimization consists of tweaking compiled C code. </I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>You can write good assembly, bad assembly, or assembly that is virtually indistinguishable from compiled code; you are more likely than not to write the latter if you think that optimization consists of tweaking compiled C code. </I></SMALL>
</TABLE>
<P>Sure, you can probably use the registers more efficiently and take advantage of an instruction or two that the compiler missed, but the code isn&#146;t going to get a whole lot faster that way.
</P>
@ -72,7 +72,7 @@ jz Match
<H3><A NAME="Heading8"></A><FONT COLOR="#000077">Taking It to the Limit</FONT></H3>
<P>The ultimate in assembly language optimization comes when you change the rules; that is, when you reorganize the entire program to allow the use of better assembly language code in the small section of code that most affects overall performance. For example, consider that the data searched in the last example is stored in an array of structures, with each structure in the array containing other information as well. In this situation, <B>REP SCASW</B> couldn&#146;t be used because the data searched through wouldn&#146;t be contiguous.</P>
<P>However, if the need for performance in searching the array is urgent enough, there&#146;s no reason why you can&#146;t reorganize the data. This might mean removing the array elements from the structures and storing them in their own array so that <B>REP SCASW</B> <I>could</I> be used.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/08-05i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Organizing a program&#146;s data so that the performance of the critical sections can be optimized is a key part of design, and one that&#146;s easily shortchanged unless, during the design stage, you thoroughly understand and work to bring together your data needs, the critical sections of your program, and potential assembly language optimizations.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Organizing a program&#146;s data so that the performance of the critical sections can be optimized is a key part of design, and one that&#146;s easily shortchanged unless, during the design stage, you thoroughly understand and work to bring together your data needs, the critical sections of your program, and potential assembly language optimizations.</I></SMALL>
</TABLE>
<P>More on this shortly.
</P>

View file

@ -78,13 +78,13 @@ ADD AX,BX ;*80
<!-- END CODE SNIP //-->
<H4 ALIGN="LEFT"><A NAME="Heading5"></A><FONT COLOR="#000077">Speeding Up Multiplication</FONT></H4>
<P>That brings us to multiplication, one of the slowest of x86 operations and one that allows for considerable optimization. One way to speed up multiplication is to use shift and add, <B>LEA</B>, or a lookup table to hard-code a multiplication operation for a fixed multiplier, as shown above. Another is to take advantage of the early-out feature of the 386 (and the 486, but in the interests of brevity I&#146;ll just say &#147;386&#148; from now on) by arranging your operands so that the multiplier (always the rightmost operand following <B>MUL</B> or <B>IMUL</B>) is no larger than the other operand.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/09-01i.jpg"><TD WIDTH="95%" VALIGN="TOP" ALIGN="LEFT"><SMALL><I>Why? Because the 386 processes one multiplier bit per cycle and immediately ends a multiplication when all significant bits of the multiplier have been processed, so fewer cycles are required to multiply a large multiplicand times a small multiplier than a small multiplicand times a large multiplier, by a factor of about 1 cycle for each significant multiplier bit eliminated.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/i.jpg"><TD WIDTH="95%" VALIGN="TOP" ALIGN="LEFT"><SMALL><I>Why? Because the 386 processes one multiplier bit per cycle and immediately ends a multiplication when all significant bits of the multiplier have been processed, so fewer cycles are required to multiply a large multiplicand times a small multiplier than a small multiplicand times a large multiplier, by a factor of about 1 cycle for each significant multiplier bit eliminated.</I></SMALL>
</TABLE>
<P>(There&#146;s a minimum execution time on this trick; below 3 significant multiplier bits, no additional cycles are saved.) For example, multiplication of 32,767 times 1 is 12 cycles faster than multiplication of 1 times 32,727.
</P>
<P>Choosing the right operand as the multiplier can work wonders. According to published specs, the 386 takes 38 cycles to multiply by a multiplier with 32 significant bits but only 9 cycles to multiply by a multiplier of 2, a performance improvement of more than four times! (My tests regularly indicate that multiplication takes 3 to 4 cycles longer than the specs indicate, but the cycle-per-bit advantage of smaller multipliers holds true nonetheless.)</P>
<P>This highlights another interesting point: <B>MUL</B> and <B>IMUL</B> on the 386 are so fast that alternative multiplication approaches, while generally still faster, are worthwhile only in truly time-critical code.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/09-02i.jpg"><TD WIDTH="95%" VALIGN="TOP" ALIGN="LEFT"><SMALL><I>On 386SXs and uncached 386s, where code size can significantly affect performance due to instruction prefetching, the compact <B>MUL</B> and <B>IMUL</B> instructions can approach and in some cases even outperform the &#147;optimized&#148; alternatives.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/i.jpg"><TD WIDTH="95%" VALIGN="TOP" ALIGN="LEFT"><SMALL><I>On 386SXs and uncached 386s, where code size can significantly affect performance due to instruction prefetching, the compact <B>MUL</B> and <B>IMUL</B> instructions can approach and in some cases even outperform the &#147;optimized&#148; alternatives.</I></SMALL>
</TABLE>
<P>All in all, <B>MUL</B> and <B>IMUL</B> are reasonable performers on the 386, no longer to be avoided in most cases&#151;and you can help that along by arranging your code to make the smaller operand the multiplier whenever you know which operand is smaller.</P>
<P>That doesn&#146;t mean that your code should test and swap operands to make sure the smaller one is the multiplier; that rarely pays off. I&#146;m speaking more of the case where you&#146;re scaling an array up by a value that&#146;s always in the range of, say, 2 to 10; because the scale value will always be small and the array elements may have any value, the scale value is the logical choice for the multiplier.</P>
@ -94,7 +94,7 @@ ADD AX,BX ;*80
<BR><A HREF="javascript:displayWindow('images/09-01.jpg',406,306)"> --><FONT COLOR="#000077"><B>Figure 9.1</B></FONT></A>&nbsp;&nbsp;<I>Simple searching method for locating a text string.</I>
</P>
<P>Rob&#146;s revelation, which he credits without explanation to Edgar Allen Poe (search nevermore?), was that by far the slowest part of the whole deal is handling <B>REPNZ SCASB</B> matches, which require checking the remainder of the string with <B>REPZ CMPS</B> and restarting <B>REPNZ SCASB</B> if no match is found.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/09-03i.jpg"><TD WIDTH="95%" VALIGN="TOP" ALIGN="LEFT"><SMALL><I>Rob points out that the number of <B>REPNZ SCASB</B> matches can easily be reduced simply by scanning for the character in the searched-for string that appears least often in the buffer being searched.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/i.jpg"><TD WIDTH="95%" VALIGN="TOP" ALIGN="LEFT"><SMALL><I>Rob points out that the number of <B>REPNZ SCASB</B> matches can easily be reduced simply by scanning for the character in the searched-for string that appears least often in the buffer being searched.</I></SMALL>
</TABLE>
<P>Imagine, if you will, that you&#146;re searching for the string &#147;EQUAL.&#148; By my approach, you&#146;d use <B>REPNZ SCASB</B> to scan for each occurrence of &#147;E,&#148; which crops up quite often in normal text. Rob points out that it would make more sense to scan for &#147;Q,&#148; then back up one character and check the whole string when a &#147;Q&#148; is found, as shown in Figure 9.2. &#147;Q&#148; is likely to occur much less often, resulting in many fewer whole-string checks and much faster processing.</P><P><BR></P>
<CENTER>

View file

@ -38,7 +38,7 @@
<P><BR></P>
<P>Listing 9.1 implements the scan-on-first-character approach. Listing 9.2 scans for whatever character the caller specifies. Listing 9.3 is a test program used to compare the two approaches. How much difference does Rob&#146;s revelation make? Plenty. Even when the entire C function call to <B>FindString</B> is timed&#151;<B>strlen</B> calls, parameter pushing, calling, setup, and all&#151;the version of <B>FindString</B> in Listing 9.2, which is directed by Listing 9.3 to scan for the infrequently-occurring &#147;Q,&#148; is about 40 percent faster on a 20 MHz cached 386 for the test search of Listing 9.3 than is the version of <B>FindString</B> in Listing 9.1, which always scans for the first character, in this case &#147;E.&#148; However, when only the search loops (the code that actually does the searching) in the two versions of <B>FindString</B> are compared, Listing 9.2 is more than <I>twice</I> as fast as Listing 9.1&#151;a remarkable improvement over code that already uses <B>REPNZ SCASB</B> and <B>REPZ CMPS</B>.</P>
<P>What I like so much about Rob&#146;s approach is that it demonstrates that optimization involves much more than instruction selection and cycle counting. Listings 9.1 and 9.2 use pretty much the same instructions, and even use the same approach of scanning with <B>REPNZ SCASB</B> and using <B>REPZ CMPS</B> to check scanning matches.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/09-04i.jpg"><TD WIDTH="95%" VALIGN="TOP" ALIGN="LEFT"><SMALL><I>The difference between Listings 9.1 and 9.2 (which gives you more than a doubling of performance) is due entirely to understanding the nature of the data being handled, and biasing the code to reflect that knowledge.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/i.jpg"><TD WIDTH="95%" VALIGN="TOP" ALIGN="LEFT"><SMALL><I>The difference between Listings 9.1 and 9.2 (which gives you more than a doubling of performance) is due entirely to understanding the nature of the data being handled, and biasing the code to reflect that knowledge.</I></SMALL>
</TABLE>
<P><A NAME="Fig2"><!-- </A><A HREF="javascript:displayWindow('images/09-02.jpg',409,306 )"> --><IMG SRC="images/09-02.jpg"><BR><!-- </A>
<BR><A HREF="javascript:displayWindow('images/09-02.jpg',409,306)"> --><FONT COLOR="#000077"><B>Figure 9.2</B></FONT></A>&nbsp;&nbsp;<I>Faster searching method for locating a text string.</I>

View file

@ -132,7 +132,7 @@ main() {
<P>Way back in Volume 1, Number 1 of <I>PC TECHNIQUES</I>, (April/May 1990) I wrote the very first of that magazine&#146;s HAX (#1), which extolled the virtues of placing your most commonly-used automatic (stack-based) variables within the stack&#146;s &#147;sweet spot,&#148; the area between &#43;127 to -128 bytes away from BP, the stack frame pointer. The reason was that the 8088 can store addressing displacements that fall within that range in a single byte; larger displacements require a full word of storage, increasing code size by a byte per instruction, and thereby slowing down performance due to increased instruction fetching time.</P>
<P>This takes on new prominence in 386 native mode, where straying from the sweet spot costs not one, but two or three bytes. Where the 8088 had two possible displacement sizes, either byte or word, on the 386 there are three possible sizes: byte, word, or dword. In native mode (32-bit protected mode), however, a prefix byte is needed in order to use a word-sized displacement, so a variable located outside the sweet spot requires either two extra bytes (an extra displacement byte plus a prefix byte) or three extra bytes (a dword displacement rather than a byte displacement). Either way, instructions grow alarmingly.</P>
<P>Performance may or may not suffer from missing the sweet spot, depending on the processor, the memory architecture, and the code mix. On a 486, prefix bytes often cost a cycle; on a 386SX, increased code size often slows performance because instructions must be fetched through the half-pint 16-bit bus; on a 386, the effect depends on the instruction mix and whether there&#146;s a cache.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/09-05i.jpg"><TD WIDTH="95%" VALIGN="TOP" ALIGN="LEFT"><SMALL><I>On balance, though, it&#146;s as important to keep your most-used variables in the stack&#146;s sweet spot in 386 native mode as it was on the 8088.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/i.jpg"><TD WIDTH="95%" VALIGN="TOP" ALIGN="LEFT"><SMALL><I>On balance, though, it&#146;s as important to keep your most-used variables in the stack&#146;s sweet spot in 386 native mode as it was on the 8088.</I></SMALL>
</TABLE>
<P>In assembly, it&#146;s easy to control the organization of your stack frame. In C, however, you&#146;ll have to figure out the allocation scheme your compiler uses to allocate automatic variables, and declare automatics appropriately to produce the desired effect. It can be done: I did it in Turbo C some years back, and trimmed the size of a program (admittedly, a large one) by several K&#151;not bad, when you consider that the &#147;sweet spot&#148; optimization is essentially free, with no code reorganization, change in logic, or heavy thinking involved.
</P><P><BR></P>

View file

@ -46,7 +46,7 @@
<P>It goes without saying that pattern matching is good; more than that, it&#146;s a large part of what we are, and, generally, the faster we are at it, the better. Not always, though. Sometimes insufficient information really is insufficient, and, in our haste to get the heady rush of coming up with a solution, incorrect or less-than-optimal conclusions are reached, as anyone who has ever done the <I>Times</I> Sunday crossword will attest. Still, my grandfather does that puzzle every Sunday <I>in ink</I>. What&#146;s his secret? Patience and discipline. He never fills a word in until he&#146;s confirmed it in his head via intersecting words, no matter how strong the urge may be to put something down where he can see it and feel like he&#146;s getting somewhere.</P>
<P>There&#146;s a surprisingly close parallel to programming here. Programming is certainly a sort of pattern matching in the sense I&#146;ve described above, and, as with crossword puzzles, following your programming instincts too quickly can be a liability. For many programmers, myself included, there&#146;s a strong urge to find a workable approach to a particular problem and start coding it <I>right now</I>, what some people call &#147;hacking&#148; a program. Going with the first thing your programming pattern matcher comes up with can be a lot of fun; there&#146;s instant gratification and a feeling of unbounded creativity. Personally, I&#146;ve always hungered to get results from my work as soon as possible; I gravitated toward graphics for its instant and very visible gratification. Over time, however, I&#146;ve learned patience.</P>
<TABLE WIDTH="100%"><TR>
<TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/10-01i.jpg"><TD WIDTH="95%"><SMALL><I>I&#146;ve come to spend an increasingly large portion of my time choosing algorithms, designing, and simply giving my mind quiet time in which to work on problems and come up with non-obvious approaches before coding; and I&#146;ve found that the extra time up front more than pays for itself in both decreased coding time and superior programs.</I></SMALL>
<TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>I&#146;ve come to spend an increasingly large portion of my time choosing algorithms, designing, and simply giving my mind quiet time in which to work on problems and come up with non-obvious approaches before coding; and I&#146;ve found that the extra time up front more than pays for itself in both decreased coding time and superior programs.</I></SMALL>
</TABLE>
<P>In this chapter, I&#146;m going to walk you through a simple but illustrative case history that nicely points up the wisdom of delaying gratification when faced with programming problems, so that your mind has time to chew on the problems from other angles. The alternative solutions you find by doing this may seem obvious, once you&#146;ve come up with them. They may not even differ greatly from your initial solutions. Often, however, they will be much better&#151;and you&#146;ll never even have the chance to decide whether they&#146;re better or not if you take the first thing that comes into your head and run with it.
</P>

View file

@ -91,7 +91,7 @@ static unsigned int gcd_recurs(unsigned int larger_int,
</PRE>
<!-- END CODE //-->
<P>As you can see from Table 10.1, Euclid&#146;s algorithm is superior, especially for large numbers (and imagine if we were working with large <I>longs!</I>).</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/10-02i.jpg"><TD WIDTH="95%"><SMALL><I>Had I been implementing GCD determination without Sedgewick&#146;s help, I would surely not have settled for Listing 10.1&#151;but I might well have ended up with Listing 10.2 in my enthusiasm over the &#147;brilliant&#148; discovery of subtracting the lesser Using Euclid&#146;s algorithm to find a GCD number from the greater. In a commercial product, my lack of patience and discipline could have been costly indeed.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>Had I been implementing GCD determination without Sedgewick&#146;s help, I would surely not have settled for Listing 10.1&#151;but I might well have ended up with Listing 10.2 in my enthusiasm over the &#147;brilliant&#148; discovery of subtracting the lesser Using Euclid&#146;s algorithm to find a GCD number from the greater. In a commercial product, my lack of patience and discipline could have been costly indeed.</I></SMALL>
</TABLE>
<P><A NAME="Fig3"><!-- </A><A HREF="javascript:displayWindow('images/10-03.jpg',411,279 )"> --><IMG SRC="images/10-03.jpg"><BR><!-- </A>
<BR><A HREF="javascript:displayWindow('images/10-03.jpg',411,279)"> --><FONT COLOR="#000077"><B>Figure 10.3</B></FONT></A>&nbsp;&nbsp;<I>Using Euclid&#146;s algorithm to find a GCD.</I>

View file

@ -129,7 +129,7 @@ _gcd endp
</PRE>
<!-- END CODE //-->
<P>Assembly language optimization is pattern matching on a local scale. Frankly, it&#146;s also the sort of boring, brute-force work that people are lousy at; compilers could out-optimize you at this level with one pass tied behind their back <I>if</I> they knew as much about the code you&#146;re writing as you do, which they don&#146;t.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/10-03i.jpg"><TD WIDTH="95%"><SMALL><I>Design optimization&#151;conceptual breakthroughs in understanding the relationships between the needs of an application, the nature of the data the application works with, and what the computer can do&#151;is global pattern matching.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>Design optimization&#151;conceptual breakthroughs in understanding the relationships between the needs of an application, the nature of the data the application works with, and what the computer can do&#151;is global pattern matching.</I></SMALL>
</TABLE>
<P>Computers are <I>much</I> worse at that sort of pattern matching than humans; computers have no way to integrate vast amounts of disparate information, much of it only vaguely defined or subject to change. People, oddly enough, are <I>better</I> at global optimization than at local optimization. For one thing, it&#146;s more interesting. For another, it&#146;s complex and imprecise enough to allow intuition and inspiration, two vastly underrated programming tools, to come to the fore. And, as I pointed out earlier, people tend to perform instantaneous solutions to even the most complex problems, while computers bog down in geometrically or exponentially increasing execution times. Oh, it may take days or weeks for a person to absorb enough information to be able to reach a solution, and the solution may only be near-optimal&#151;but the solution itself (or, at least, each of the pieces of the solution) arrives in a flash.</P>
<P>Those flashes are your programming pattern matcher doing its job. <I>Your</I> job is to give your pattern matcher the opportunity to get to know each problem and run through it two or three times, from different angles, to see what unexpected solutions it can come up with.</P>

View file

@ -45,7 +45,7 @@
<H4 ALIGN="CENTER"><A NAME="Heading6"></A><FONT COLOR="#000077">System Wait States</FONT></H4>
<P>The 286 and 386 were designed to lose relatively little performance to the prefetch queue cycle-eater...<I>when used with zero-wait-state memory:</I> memory that can complete memory accesses so rapidly that no wait states are needed. However, true zero-wait-state memory is almost never used with those processors. Why? Because memory that can keep up with a 286 is fairly expensive, and memory that can keep up with a 386 is <I>very</I> expensive. Instead, computer designers use alternative memory architectures that offer more performance for the dollar&#151;but less performance overall&#151;than zero-wait-state memory. (It <I>is</I> possible to build zero-wait-state systems for the 286 and 386; it&#146;s just so expensive that it&#146;s rarely done.)</P>
<P>The IBM AT and true compatibles use one-wait-state memory (some AT clones use zero-wait-state memory, but such clones are less common than one-wait-state AT clones). The 386 systems use a wide variety of memory systems&#151;including high-speed caches, interleaved memory, and static-column RAM&#151;that insert anywhere from 0 to about 5 wait states (and many more if 8 or 16-bit memory expansion cards are used); the exact number of wait states inserted at any given time depends on the interaction between the code being executed and the memory system it&#146;s running on.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/11-01i.jpg"><TD WIDTH="95%" ALIGN="LEFT" VALIGN="TOP"><SMALL><I>The performance of most 386 memory systems can vary greatly from one memory access to another, depending on factors such as what data happens to be in the cache and which interleaved bank and/or RAM column was accessed last.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/i.jpg"><TD WIDTH="95%" ALIGN="LEFT" VALIGN="TOP"><SMALL><I>The performance of most 386 memory systems can vary greatly from one memory access to another, depending on factors such as what data happens to be in the cache and which interleaved bank and/or RAM column was accessed last.</I></SMALL>
</TABLE>
<P>The many memory systems in use make it impossible for us to optimize for 286/386 computers with the precision that&#146;s possible on the 8088. Instead, we must write code that runs reasonably well under the varying conditions found in the 286/386 arena.
</P>

View file

@ -60,7 +60,7 @@
</P>
<P>Figure 11.1 illustrates this phenomenon. The conversion of word-sized accesses to odd addresses into double byte-sized accesses is transparent to memory-accessing instructions; all any instruction knows is that the requested word has been accessed, no matter whether 1 word-sized access or 2 byte-sized accesses were required to accomplish it.</P>
<P>The penalty for performing a word-sized access starting at an odd address is easy to calculate: Two accesses take twice as long as one access.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/11-02i.jpg"><TD WIDTH="95%" ALIGN="LEFT" VALIGN="TOP"><SMALL><I>In other words, the effective capacity of the 286&#146;s external data bus is </I><I>halved</I> <I>when a word-sized access to an odd address is performed.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/i.jpg"><TD WIDTH="95%" ALIGN="LEFT" VALIGN="TOP"><SMALL><I>In other words, the effective capacity of the 286&#146;s external data bus is </I><I>halved</I> <I>when a word-sized access to an odd address is performed.</I></SMALL>
</TABLE>
<P>That, in a nutshell, is the data alignment cycle-eater, the one new cycle-eater of the 286 and 386. (The data alignment cycle-eater is a close relative of the 8088&#146;s 8-bit bus cycle-eater, but since it behaves differently&#151;occurring only at odd addresses&#151;and is avoided with a different workaround, we&#146;ll consider it to be a new cycle-eater.)
</P>

View file

@ -67,7 +67,7 @@ LoopTop:
</PRE>
<!-- END CODE SNIP //-->
<P>While word-aligning branch destinations can improve branching performance, it&#146;s a nuisance and can increase code size a good deal, so it&#146;s not worth doing in most code. Besides, <B>EVEN</B> inserts a <B>NOP</B> instruction if necessary, and the time required to execute a <B>NOP</B> can sometimes cancel the performance advantage of having a word-aligned branch destination.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/11-03i.jpg"><TD WIDTH="95%" VALIGN="TOP" ALIGN="LEFT"><SMALL><I>Consequently, it&#146;s best to word-align only those branch destinations that can be reached solely by branching.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/i.jpg"><TD WIDTH="95%" VALIGN="TOP" ALIGN="LEFT"><SMALL><I>Consequently, it&#146;s best to word-align only those branch destinations that can be reached solely by branching.</I></SMALL>
</TABLE>
<P>I recommend that you only go out of your way to word-align the start offsets of your subroutines, as in:
</P>
@ -88,7 +88,7 @@ FindChar proc near
<H4 ALIGN="CENTER"><A NAME="Heading10"></A><FONT COLOR="#000077">Alignment and the Stack</FONT></H4>
<P>One side-effect of the data alignment cycle-eater of the 286 and 386 is that you should <I>never</I> allow the stack pointer to become odd. (You can make the stack pointer odd by adding an odd value to it or subtracting an odd value from it, or by loading it with an odd value.) An odd stack pointer on the 286 or 386 (or a non-doubleword-aligned stack in 32-bit protected mode on the 386, 486, or Pentium) will significantly reduce the performance of <B>PUSH,</B> <B>POP,</B> <B>CALL</B>, and <B>RET</B>, as well as <B>INT</B> and <B>IRET</B>, which are executed to invoke DOS and BIOS functions, handle keystrokes and incoming serial characters, and manage the mouse. I know of a Forth programmer who vastly improved the performance of a complex application on the AT simply by forcing the Forth interpreter to maintain an even stack pointer at all times.</P>
<P>An interesting corollary to this rule is that you shouldn&#146;t <B>INC SP</B> twice to add 2, even though that takes fewer bytes than <B>ADD SP,2</B>. The stack pointer is odd between the first and second <B>INC</B>, so any interrupt occurring between the two instructions will be serviced more slowly than it normally would. The same goes for decrementing twice; use <B>SUB SP,2</B> instead.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/11-04i.jpg"><TD WIDTH="95%" ALIGN="LEFT" VALIGN="TOP"><SMALL><I>Keep the stack pointer aligned at all times.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/i.jpg"><TD WIDTH="95%" ALIGN="LEFT" VALIGN="TOP"><SMALL><I>Keep the stack pointer aligned at all times.</I></SMALL>
</TABLE>
<H4 ALIGN="CENTER"><A NAME="Heading11"></A><FONT COLOR="#000077">The DRAM Refresh Cycle-Eater: Still an Act of God</FONT></H4>
<P>The DRAM refresh cycle-eater is the cycle-eater that&#146;s least changed from its 8088 form on the 286 and 386. In the AT, DRAM refresh uses a little over five percent of all available memory accesses, slightly less than it uses in the PC, but in the same ballpark. While the DRAM refresh penalty varies somewhat on various AT clones and 386 computers (in fact, a few computers are built around static RAM, which requires no refresh at all; likewise, caches are made of static RAM so cached systems generally suffer less from DRAM refresh), the 5 percent figure is a good rule of thumb.

View file

@ -42,7 +42,7 @@
<P>In other words, instructions that access display memory won&#146;t run a whole lot faster on ATs and faster computers than they do on PCs. That explains one of the two viewpoints expressed at the beginning of this section: The display adapter cycle-eater is just about the same on high-end computers as it is on the PC, in the sense that it allows instructions that access display memory to run at just about the same speed on all computers.</P>
<P>Of course, the picture is quite a bit different when you compare the performance of instructions that access display memory to the <I>maximum</I> performance of those instructions. Instructions that access display memory receive many more wait states when running on a 286 than they do on an 8088. Why? While the 286 is capable of accessing memory much more often than the 8088, we&#146;ve seen that the frequency of access to display memory is determined not by processor speed but by the display adapter itself. As a result, both processors are actually allowed just about the same maximum number of accesses to display memory in any given time. By definition, then, the 286 must spend many more cycles waiting than does the 8088.</P>
<P>And that explains the second viewpoint expressed above regarding the display adapter cycle-eater vis-a-vis the 286 and 386. The display adapter cycle-eater, as measured in cycles lost to wait states, is indeed much worse on AT-class computers than it is on the PC, and it&#146;s worse still on more powerful computers.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/11-05i.jpg"><TD WIDTH="95%" VALIGN="TOP" ALIGN="LEFT"><SMALL><I>How bad is the display adapter cycle-eater on an AT? It&#146;s this bad: Based on my (not inconsiderable) experience in timing display adapter access, I&#146;ve found that the display adapter cycle-eater can slow an AT&#151;or even a 386 computer&#151;to near-PC speeds when display memory is accessed.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/i.jpg"><TD WIDTH="95%" VALIGN="TOP" ALIGN="LEFT"><SMALL><I>How bad is the display adapter cycle-eater on an AT? It&#146;s this bad: Based on my (not inconsiderable) experience in timing display adapter access, I&#146;ve found that the display adapter cycle-eater can slow an AT&#151;or even a 386 computer&#151;to near-PC speeds when display memory is accessed.</I></SMALL>
</TABLE>
<P>I know that&#146;s hard to believe, but the display adapter cycle-eater gives out just so many display memory accesses in a given time, and no more, no matter how fast the processor is. In fact, the faster the processor, the more the display adapter cycle-eater hurts the performance of instructions that access display memory. The display adapter cycle-eater is not only still present in 286/386 computers, it&#146;s worse than ever.
</P>

View file

@ -37,7 +37,7 @@
</CENTER>
<P><BR></P>
<P>Finally, it&#146;s possible in real mode to use the 386&#146;s new addressing modes, in which <I>any</I> 32-bit general-purpose register or pair of registers can be used to address memory. What&#146;s more, multiplication of memory-addressing registers by 2, 4, or 8 for look-ups in word, doubleword, or quadword tables can be built right into the memory addressing mode. (The 32-bit addressing modes are discussed further in later chapters.) In protected mode, these new addressing modes allow you to address a full 4 gigabytes per segment, but in real mode you&#146;re still limited to 64K, even with 32-bit registers and the new addressing modes, unless you play some unorthodox tricks with the segment registers.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/11-06i.jpg"><TD WIDTH="95%" ALIGN="LEFT" VALIGN="TOP"><SMALL><I>Note well: Those tricks don&#146;t necessarily work with system software such as Windows, so I&#146;d recommend against using them. If you want 4-gigabyte segments, use a 32-bit environment such as Win32.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/i.jpg"><TD WIDTH="95%" ALIGN="LEFT" VALIGN="TOP"><SMALL><I>Note well: Those tricks don&#146;t necessarily work with system software such as Windows, so I&#146;d recommend against using them. If you want 4-gigabyte segments, use a 32-bit environment such as Win32.</I></SMALL>
</TABLE>
<H4 ALIGN="CENTER"><A NAME="Heading15"></A><FONT COLOR="#000077">Optimization Rules: The More Things Change...</FONT></H4>
<P>Let&#146;s see what we&#146;ve learned about 286/386 optimization. Mostly what we&#146;ve learned is that our familiar PC cycle-eaters still apply, although in somewhat different forms, and that the major optimization rules for the PC hold true on ATs and 386-based computers. You won&#146;t go wrong on any of these computers if you keep your instructions short, use the registers heavily and avoid memory, don&#146;t branch, and avoid accessing display memory like the plague.

View file

@ -50,7 +50,7 @@
<P>An encyclopedic approach to 486 optimization would take a book all by itself, so in this chapter I&#146;m only going to hit the highlights of 486 optimization, touching on several optimization rules, some documented, some not. You might also want to check out the following sources of 486 information: <I>i486 Microprocessor Programmer&#146;s Reference Manual,</I> from Intel; &#147;8086 Optimization: Aim Down the Middle and Pray,&#148; in the March, 1991 <I>Dr. Dobb&#146;s Journal</I>; and &#147;Peak Performance: On to the 486,&#148; in the November, 1990 <I>Programmer&#146;s Journal.</I></P>
<H3><A NAME="Heading4"></A><FONT COLOR="#000077">Rules to Optimize By</FONT></H3>
<P>In Appendix G of the <I>i486 Microprocessor Programmer</I>&#146;<I>s</I> <I>Reference Manual</I>, Intel lists a number of optimization techniques for the 486. While neither exhaustive (we&#146;ll look at two undocumented optimizations shortly) nor entirely accurate (we&#146;ll correct two of the rules here), Intel&#146;s list is certainly a good starting point. In particular, the list conveys the extent to which 486 optimization differs from optimization for earlier x86 processors. Generally, I&#146;ll be discussing optimization for real mode (it being the most widely used mode at the moment), although many of the rules should apply to protected mode as well.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/12-01i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>486 optimization is generally more precise and less frustrating than optimization for other x86 processors because every 486 has an identical internal cache. Whenever both the instructions being executed and the data the instructions access are in the cache, those instructions will run in a consistent and calculatable number of cycles on all 486s, with little chance of interference from the prefetch queue and without regard to the speed of external memory.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>486 optimization is generally more precise and less frustrating than optimization for other x86 processors because every 486 has an identical internal cache. Whenever both the instructions being executed and the data the instructions access are in the cache, those instructions will run in a consistent and calculatable number of cycles on all 486s, with little chance of interference from the prefetch queue and without regard to the speed of external memory.</I></SMALL>
</TABLE>
<P>In other words, for cached code (which time-critical code almost always is), performance is predictable and can be calculated with good precision, and those calculations will apply on any 486. However, &#147;predictable&#148; doesn&#146;t mean &#147;trivial&#148;; the cycle times printed for the various instructions are not the whole story. You must be aware of all the rules, documented and undocumented, that go into calculating actual execution times&#151;and uncovering some of those rules is exactly what this chapter is about.
</P>
@ -59,7 +59,7 @@
</P>
<P>Intel cautions against using indexing to address memory because there&#146;s a one-cycle penalty for indexed addressing. True enough&#151;but &#147;indexed addressing&#148; might not mean what you expect.</P>
<P>Traditionally, SI and DI are considered the index registers of the x86 CPUs. That is not the sense in which &#147;indexed addressing&#148; is meant here, however. In real mode, indexed addressing means that two registers, rather than one or none, are used to point to memory. (In this context, the use of one register to address memory is &#147;base addressing,&#148; no matter what register is used.) <B>MOV AX, [BX&#43;DI]</B> and <B>MOV CL, [BP&#43;SI&#43;10]</B> perform indexed addressing; <B>MOV AX,[BX]</B> and <B>MOV DL, [SI&#43;1]</B> do not.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/12-02i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Therefore, in real mode, the rule is to avoid using two registers to point to memory whenever possible. Often, this simply means adding the two registers together outside a loop before memory is actually addressed.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Therefore, in real mode, the rule is to avoid using two registers to point to memory whenever possible. Often, this simply means adding the two registers together outside a loop before memory is actually addressed.</I></SMALL>
</TABLE>
<P>As an example, you might adhere to this rule by replacing the code
</P>

View file

@ -76,7 +76,7 @@ LoopTop:
<!-- END CODE SNIP //-->
<P>Now that we understand what Intel means by this rule, let me make a very important comment: My observations indicate that for real-mode code, the documentation understates the extent of the penalty for interrupting the address calculation pipeline by loading a memory pointer just before it&#146;s used.
</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/12-03i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>The truth of the matter appears to be that if a register is the destination of one instruction and is then used by the next instruction to address memory in real mode, not one but two cycles are lost!</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>The truth of the matter appears to be that if a register is the destination of one instruction and is then used by the next instruction to address memory in real mode, not one but two cycles are lost!</I></SMALL>
</TABLE>
<P>In 32-bit protected mode, however, the penalty is, in fact, the 1 cycle that Intel .
</P>

View file

@ -118,7 +118,7 @@ xlat
</PRE>
<!-- END CODE SNIP //-->
<P>even though AL must be converted to a word by <B>XLAT</B> before it can be added to BX and used to address memory. In fact, none of the penalties mentioned in this chapter apply to <B>XLAT</B>, apparently because <B>XLAT</B> is so slow&#151;4 cycles&#151;that it gives the 486 time to perform addressing calculations during the course of the instruction.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/12-04i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>While it&#146;s nice that <B>XLAT</B> doesn&#146;t suffer from the various 486 addressing penalties, the reason for that is basically that <B>XLAT</B> is slow, so there&#146;s still no compelling reason to use <B>XLAT</B> on the 486.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>While it&#146;s nice that <B>XLAT</B> doesn&#146;t suffer from the various 486 addressing penalties, the reason for that is basically that <B>XLAT</B> is slow, so there&#146;s still no compelling reason to use <B>XLAT</B> on the 486.</I></SMALL>
</TABLE>
<P>In general, penalties for interrupting the 486&#146;s pipeline apply primarily to the fast core instructions of the 486, most notably register-only instructions and <B>MOV</B>, although arithmetic and logical operations that access memory are also often affected. I don&#146;t know all the performance dependencies, and I don&#146;t plan to; figuring all of them out would be a big, boring job of little value. Basically, on the 486 you should concentrate on using those fast core instructions when performance matters, and all the rules I&#146;ll discuss do indeed apply to those instructions.</P><P><BR></P>
<CENTER>

View file

@ -76,7 +76,7 @@ mov ax,[bx]
</PRE>
<!-- END CODE SNIP //-->
<P>A more sophisticated programmer would expect to lose one cycle, because BX is loaded two cycles before being used to address memory. In fact, though, this code takes 5 cycles&#151;2 cycles, or 67 percent, longer than normal. Why? Well, under normal conditions, loading a byte register&#151;CL in this case&#151;one cycle before using a register to address memory produces no penalty; loading 2 cycles ahead is the only case that normally incurs a penalty. However, think of Rule #4 as meaning that loading a byte register disrupts the memory addressing pipeline as it starts up. Viewed that way, we can see that <B>MOV BX,OFFSET MemVar</B> interrupts the addressing pipeline, forcing it to start again, and then, presumably, <B>MOV CL,AL</B> interrupts the pipeline again because the pipeline is now on its first cycle: the one that loading a byte register can affect.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/12-05i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>I know&#151;it seems awfully complicated. It isn&#146;t, really. Generally, try not to use byte destinations exactly two cycles before using a register to address memory, and try not to load a register either one or two cycles before using it to address memory, and you&#146;ll be fine.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>I know&#151;it seems awfully complicated. It isn&#146;t, really. Generally, try not to use byte destinations exactly two cycles before using a register to address memory, and try not to load a register either one or two cycles before using it to address memory, and you&#146;ll be fine.</I></SMALL>
</TABLE>
<H4 ALIGN="LEFT"><A NAME="Heading11"></A><FONT COLOR="#000077">Timing Your Own 486 Code</FONT></H4>
<P>In case you want to do some 486 performance analysis of your own, let me show you how I arrived at one of the above conclusions; at the same time, I can warn you of the timing hazards of the cache. Listings 12.1 and 12.2 show the code I ran through the Zen timer in order to establish the effects of loading a byte register before using a register to address memory. Listing 12.1 ran in 120 &#181;s on a 33 MHz 486, or 4 cycles per repetition (120 &#181;s/1000 repetitions = 120 ns per repetition; 120 ns per repetition/30 ns per cycle = 4 cycles per repetition); Listing 12.2 ran in 90 &#181;s, or 3 cycles, establishing that loading a byte register costs a cycle only when it&#146;s performed exactly 2 cycles before addressing memory.
@ -126,7 +126,7 @@ Done:
</PRE>
<!-- END CODE //-->
<P>Note that Listings 12.1 and 12.2 each repeat the timing of the code under test a second time, to make sure that the instructions are in the cache on the second pass, the one for which results are displayed. Also note that the code is less than 8K in size, so that it can all fit in the 486&#146;s 8K internal cache. If I double the <B>REPT</B> value in Listing 12.2 to 2,000, making the test code larger than 8K, the execution time more than doubles to 224 &#181;s, or 3.7 cycles per repetition; the extra seven-tenths of a cycle comes from fetching non-cached instruction bytes.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/12-06i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Whenever you see non-integral timing results of this sort, it&#146;s a good bet that the test code or data isn&#146;t cached.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Whenever you see non-integral timing results of this sort, it&#146;s a good bet that the test code or data isn&#146;t cached.</I></SMALL>
</TABLE>
<H3><A NAME="Heading12"></A><FONT COLOR="#000077">The Story Continues</FONT></H3>
<P>There&#146;s certainly plenty more 486 lore to explore, including the 486&#146;s unique prefetch queue, more optimization rules, branching optimizations, performance implications of the cache, the cost of cache misses for reads, and the implications of cache write-through for writes. Nonetheless, we&#146;ve covered quite a bit of ground in this chapter, and I trust you&#146;ve gotten a feel for the considerable extent to which 486 optimization differs from what you&#146;re used to. Odd as 486 optimization is, though, it&#146;s well worth mastering, for the 486 is, at its best, so staggeringly fast that carefully crafted 486 code can do more than twice as much per cycle as the best 386 code&#151;which makes it perhaps 50 times as fast as optimized code for the original PC.

View file

@ -61,7 +61,7 @@ add dx,[bx&#43;8000h] ;increment word and line count
<BR><A HREF="javascript:displayWindow('images/13-01.jpg',100,65)"> --><FONT COLOR="#000077"><B>Figure 13.1</B></FONT></A>&nbsp;&nbsp;<I>Cycle-eaters in the original WC.</I>
</P>
<TABLE WIDTH="100%">
<TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/13-01i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Remember, pipeline penalties diminish with increasing number of cycles, not instructions, between the pipeline disrupter and the potentially affected instruction.</I></SMALL>
<TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Remember, pipeline penalties diminish with increasing number of cycles, not instructions, between the pipeline disrupter and the potentially affected instruction.</I></SMALL>
</TABLE>
<P><B>LISTING 13.2 L13-2.ASM</B></P>
<!-- CODE SNIP //-->

View file

@ -72,7 +72,7 @@ push word ptr [bx]
</P>
<P>Likewise, popping a memory location takes six cycles, but popping a register and writing it to memory takes only two cycles combined. The <I>i486 Microprocessor Programmer&#146;s Reference Manual</I> lists a 4-cycle execution time for popping a register, but pay that no mind; popping a register takes only 1 cycle.</P>
<P>Why is it that such a convenient operation as pushing or popping memory is so slow? The rule on the 486 is that simple operations, which can be executed in a single cycle by the 486&#146;s RISC core, are fast; whereas complex operations, which must be carried out in microcode just as they were on the 386, are almost all relatively slow. Slow, complex operations include all the string instructions except <B>REP MOVS,</B> as well as <B>XLAT, LOOP,</B> and, of course, <B>PUSH <I>mem</I></B> and <B>POP <I>mem.</I></B></P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/13-02i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Whenever possible, try to use the 486&#146;s 1-cycle instructions, including <B>MOV, ADD, SUB, CMP, ADC, SBB, XOR, AND, OR, TEST, LEA</B>, and <B>PUSH reg</B> and <B>POP reg</B>. These instructions have an added benefit in that it&#146;s often possible to rearrange them for maximum pipeline efficiency, as is the case with Terje&#146;s optimization described earlier in this chapter.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Whenever possible, try to use the 486&#146;s 1-cycle instructions, including <B>MOV, ADD, SUB, CMP, ADC, SBB, XOR, AND, OR, TEST, LEA</B>, and <B>PUSH reg</B> and <B>POP reg</B>. These instructions have an added benefit in that it&#146;s often possible to rearrange them for maximum pipeline efficiency, as is the case with Terje&#146;s optimization described earlier in this chapter.</I></SMALL>
</TABLE>
<H3><A NAME="Heading6"></A><FONT COLOR="#000077">Optimal 1-Bit Shifts and Rotates</FONT></H3>
<P>On a 486, the n-bit forms of the shift and rotate instructions&#151;as in <B>ROR AX,2</B> and <B>SHL BX,9</B>&#151;are 2-cycle instructions, but the 1-bit forms&#151;as in <B>ROR AX,1</B> and <B>SHL BX,1&#151;</B>are <I>3-cycle</I> instructions. Go figure.</P>

View file

@ -48,7 +48,7 @@ mov al,BaseTable[ecx&#43;edx*4]
<P>Any register may serve as the base register component of an address. Any register except ESP may also serve as the index register, which can be scaled by 1, 2, 4, or 8. (Scaling is very handy for performing lookups in arrays and tables.) The same register may serve as both base and index register, except for ESP, which can only be the base. Incidentally, it makes sense that ESP can&#146;t be scaled; ESP presumably always points to a valid stack, and I can&#146;t think of any reason you&#146;d want to use the stack pointer times 2, 4, or 8 in an address. ESP is, by its nature, a base rather than index pointer.</P>
<P>That&#146;s all there is to the functionality of 32-bit addressing; it&#146;s very simple, much simpler than 16-bit addressing, with its sharply limited memory addressing register combinations. The costs of 32-bit addressing are a bit more subtle. The only performance cost (apart from the aforementioned 1-cycle penalty for using 32-bit addressing in real mode) is a 1-cycle penalty imposed for using an index register. In this context, you use an index register when you use a register that&#146;s scaled, or when you use the sum of two registers to point to memory. <B>MOV BL,[EBX*2]</B> uses an index register and takes an extra cycle, as does <B>MOV CL,[EAX&#43;EDX]; MOV CL,[EAX&#43;100H]</B> is not indexed, however.</P>
<P>The other cost of 32-bit addressing is in instruction size. Old-style 16-bit addressing usually (except in a few special cases) uses one extra byte, which Intel calls the Mod-R/M byte, which is placed immediately after each instruction&#146;s opcode to describe the memory addressing mode, plus 1 or 2 optional bytes of addressing displacement&#151;that is, a constant value to add into the address. In many cases, 32-bit addressing continues to use the Mod-R/M byte, albeit with a different interpretation; in these cases, 32-bit addressing is no larger than 16-bit addressing, except when a 32-bit displacement is involved. For example, <B>MOV AL, [EBX]</B> is a 2-byte instruction; <B>MOV AL, [EBX&#43;10H]</B> is a 3-byte instruction; and <B>MOV AL, [EBX&#43;10000H]</B> is a 6-byte instruction.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/13-03i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Note that 1 and 4-byte displacements, but not 2-byte displacements, are supported for 32-bit addressing. Code size can be greatly improved by keeping stack frame variables within 128 bytes of EBP, and variables in pointed-to structures within 127 bytes of the start of the structure, so that displacements can be 1 rather than 4 bytes.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Note that 1 and 4-byte displacements, but not 2-byte displacements, are supported for 32-bit addressing. Code size can be greatly improved by keeping stack frame variables within 128 bytes of EBP, and variables in pointed-to structures within 127 bytes of the start of the structure, so that displacements can be 1 rather than 4 bytes.</I></SMALL>
</TABLE>
<P>However, because 32-bit addressing supports many more addressing combinations than 16-bit addressing, the Mod-R/M byte can&#146;t describe all the combinations. Therefore, whenever an index register (as described above) is involved, a second byte, the SIB byte, follows the Mod-R/M byte to provide additional address information. Consequently, whenever you use a scaled memory addressing register or use the sum of two registers to point to memory, you automatically add 1 cycle and 1 byte to that instruction. This is not to say that you shouldn&#146;t use index registers when they&#146;re needed, but if you find yourself using them inside key loops, you should see if it&#146;s possible to move the index calculation outside the loop as, for example, in a loop like this:
</P>

View file

@ -52,7 +52,7 @@
<P>String searching is the simple matter of finding the first occurrence of a particular sequence of bytes (the pattern) within another sequence of bytes (the buffer). The obvious, brute-force approach is to try every possible match location, starting at the beginning of the buffer and advancing one position after each mismatch, until either a match is found or the buffer is exhausted. There&#146;s even a nifty string instruction, <B>REPZ CMPS,</B> that&#146;s perfect for comparing the pattern to the contents of the buffer at each location. What could be simpler?</P>
<P>We have some important information that we&#146;re not yet using, though. Typically, the buffer will contain a wide variety of bytes. Let&#146;s assume that the buffer contains text, in which case there will be dozens of different characters; and although the distribution of characters won&#146;t usually be even, neither will any one character constitute half the buffer, or anything close. A reasonable conclusion is that the first character of the pattern will rarely match the first character of the buffer location currently being checked. This allows us to use the speedy <B>REPNZ SCASB</B> to whiz through the buffer, eliminating most potential match locations with single repetitions of <B>SCASB.</B> Only when that first character does (infrequently) match must we drop back to the slower <B>REPZ CMPS</B> approach.</P>
<P>It&#146;s important to understand that we&#146;re assuming that the buffer is typical text. That&#146;s what I meant at the outset, when I said that the information you need may be under your nose.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/14-01i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Formally, you don&#146;t know a blessed thing about the search buffer, but experience, common sense, and your knowledge of the application give you a great deal of useful, if somewhat imprecise, information.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Formally, you don&#146;t know a blessed thing about the search buffer, but experience, common sense, and your knowledge of the application give you a great deal of useful, if somewhat imprecise, information.</I></SMALL>
</TABLE>
<P>If the buffer contains the letter &#145;A&#146; repeated 1,000 times, followed by the letter &#145;B,&#146; then the <B>REPNZ SCASB/REPZ CMPS</B> approach will be much slower than the brute-force <B>REPZ CMPS</B> approach when searching for the pattern &#147;AB,&#148; because <B>REPNZ SCASB</B> would match at every buffer location. You could construct a horrendous worst-case scenario for almost any good optimization; the key is understanding the usual conditions under which your code will work.</P>
<P>As discussed in Chapter 9, we also know that certain characters have lower probabilities of matching than others. In a normal buffer, &#145;T&#146; will match far more often than &#145;X.&#146; Therefore, if we use <B>REPNZ SCASB</B> to scan for the least common letter in the search string, rather than the first letter, we&#146;ll greatly decrease the number of times we have to drop back to <B>REPZ CMPS,</B> and the search time will become very close to the time it takes <B>REPNZ SCASB</B> to go from the start of the buffer to the match location. If the distance to the first match is N bytes, the least-common <B>REPNZ SCASB</B> approach will take about as long as N repetitions of <B>REPNZ SCASB.</B></P>

View file

@ -41,7 +41,7 @@
<P>If that makes your head hurt, it should&#151;and don&#146;t worry. This line of thinking, which is the basis of the Knuth-Morris-Pratt algorithm and half the basis of the Boyer-Moore algorithm, is what gives Boyer-Moore its reputation for inscrutability. That reputation is well deserved for this aspect (which I will not discuss further in this book), but there&#146;s another part of Boyer-Moore that&#146;s easily understood, easily implemented, and highly effective.</P>
<P>Consider this: We&#146;re searching for the pattern &#147;ABC,&#148; beginning the search at the start (offset 0) of a buffer containing &#147;ABZABC.&#148; We match on &#145;A,&#146; we match on &#145;B,&#146; and we mismatch on &#145;C&#146;; the buffer contains a &#145;Z&#146; in this position. What have we learned? Why, we&#146;ve learned not only that the pattern doesn&#146;t match the buffer starting at offset 0, but also that it can&#146;t possibly match starting at offset 1 or offset 2, either! After all, there&#146;s a &#145;Z&#146; in the buffer at offset 2; since the pattern doesn&#146;t contain a single &#145;Z,&#146; there&#146;s no way that the pattern can match starting at <I>any</I> location from which it would span the &#145;Z&#146; at offset 2. We can just skip straight from offset 0 to offset 3 and continue, saving ourselves two comparisons.</P>
<P>Unfortunately, this approach only pays off big when a near-complete partial match is found; if the comparison fails on the first pattern character, as often happens, we can only skip ahead 1 byte, as usual. Look at it differently, though: What if we compare the pattern starting with the last (rightmost) byte, rather than the first (leftmost) byte? In other words, what if we compare from high memory toward low, in the direction in which string instructions go after the <B>STD</B> instruction? After all, we&#146;re comparing one set of bytes (the pattern) to another set of bytes (a portion of the buffer); it doesn&#146;t matter in the least in what order we compare them, so long as all the bytes in one set are compared to the corresponding bytes in the other set.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/14-02i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Why on earth would we want to start with the rightmost character? Because a mismatch on the rightmost character tells us a great deal more than a mismatch on the leftmost character.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Why on earth would we want to start with the rightmost character? Because a mismatch on the rightmost character tells us a great deal more than a mismatch on the leftmost character.</I></SMALL>
</TABLE>
<P>We learn nothing new from a mismatch on the leftmost character, except that the pattern can&#146;t match starting at that location. A mismatch on the rightmost character, however, tells us about the possibilities of the pattern matching starting at every buffer location from which the pattern spans the mismatch location. If the mismatched character in the buffer doesn&#146;t appear in the pattern, then we&#146;ve just eliminated not one potential match, but as many potential matches as there are characters in the pattern; that&#146;s how many locations there are in the buffer that <I>might</I> have matched, but have just been shown not to, because they overlap the mismatched character that doesn&#146;t belong in the pattern. In this case, we can skip ahead by the full pattern length in the buffer! This is how we can outperform even <B>REPNZ SCASB; REPNZ SCASB</B> has to check every byte in the buffer, but Boyer-Moore doesn&#146;t.</P>
<P>Figure 14.1 illustrates the operation of a Boyer-Moore search when the rightcharacter of the search pattern (which is the first character that&#146;s compared at each location because we&#146;re comparing backwards) mismatches with a buffer character that appears nowhere in the pattern. Figure 14.2 illustrates the operation of a partial match when the mismatch occurs with a character that&#146;s not a pattern member. In this case, we can only skip ahead past the mismatch location, resulting in an advance of fewer bytes than the pattern length, and potentially as little as the same single byte distance by which the standard search approach advances.</P>

View file

@ -106,7 +106,7 @@ struct LinkNode *DeleteNodeAfter(struct LinkNode **HeadOfListPtr,
<BR><A HREF="javascript:displayWindow('images/15-02.jpg',407,102)"> --><FONT COLOR="#000077"><B>Figure 15.2</B></FONT></A>&nbsp;&nbsp;<I>Using a dummy head and tail node with a linked list.</I>
</P>
<TABLE WIDTH="100%">
<TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/15-01i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>The next-node pointer of the head node, which points to the first real node, is the only part of the head node that&#146;s actually used. This way the same code works on the head node as on the rest of the list, so there are no special cases.</I></SMALL>
<TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>The next-node pointer of the head node, which points to the first real node, is the only part of the head node that&#146;s actually used. This way the same code works on the head node as on the rest of the list, so there are no special cases.</I></SMALL>
</TABLE>
<P>Likewise, there should be a separate node for the tail of the list, so that every node that contains real data is guaranteed to have a node on either side of it. In this scheme, an empty list contains two nodes, as shown in Figure 15.3. Although it is not necessary, the tail node may point to itself as its own next node, rather than contain a <B>NULL</B> pointer. This way, a deletion operation on an empty list will have no effect&#151;quite unlike the same operation performed on a list terminated with a <B>NULL</B> pointer. The tail node of a list terminated like this can be detected because it will be the only node for which the next-node pointer equals the current-node pointer.</P>
<P>Figure 15.3 is a giant step in the right direction, but we can still make a few refinements. The inner loop of any code that scans through such a list has to perform a special test on each node to determine whether the tail has been reached. So, for example, code to find the first node containing a value field greater than or equal to a certain value has to perform two tests in the inner loop, as shown in Listing 15.4.</P>

View file

@ -135,7 +135,7 @@
<P>Listing 16.4 features several interesting tricks. First, it uses <B>LODSB</B> and <B>XLAT</B> in succession, a very neat way to get a pointed-to byte, advance the pointer, and look up the value indexed by the byte in a table, all with just two instruction bytes. (Interestingly, Listing 16.4 would probably run quite a bit better still on an 8088, where <B>LODSB</B> and <B>XLAT</B> have a greater advantage over conventional instructions. On the 486 and Pentium, however, <B>LODSB</B> and <B>XLAT</B> lose much of their appeal, and should be replaced with <B>MOV</B> instructions.) Better yet, <B>LODSB</B> and <B>XLAT</B> don&#146;t alter the flags, so the Zero flag status set before <B>LODSB</B> is still around to be tested after <B>XLAT</B> .</P>
<P>Finally, if you look closely, you will see that Listing 16.4 jumps out of the loop to increment the word count in the case where a word is actually found, with a duplicate of the loop-bottom code placed after the code that increments the word count, to avoid an extra branch back into the loop; this replaces the more intuitive approach of jumping around the incrementing code to the loop bottom when a word isn&#146;t found. Although this incurs a branch every time a word is found, a word is typically found only once every 5 or 6 bytes; on average, then, a branch is saved about two-thirds of the time. This is an excellent example of how understanding the nature of the data you&#146;re processing allows you to optimize in ways the compiler can&#146;t. <I>Know your data!</I></P>
<P>So, gosh, Listing 16.4 is the best word-counting code in the universe, right? Not hardly. If there&#146;s one thing my years of toil in this vale of silicon have taught me, it&#146;s that there&#146;s never a lack of potential for further optimization. <I>Never!</I> Off the top of my head, I can think of at least three ways to speed up Listing 16.4; and, since Turbo Profiler reports that even in Listing 16.4, 88 percent of the time is spent scanning the buffer (as opposed to reading the file), there&#146;s potential for those further optimizations to improve performance significantly. (However, it is true that when access is performed to a hard rather than RAM disk, disk access jumps to about half of overall execution time.) One possible optimization is unrolling the loop, although that is truly a last resort because it tends to make further changes extremely difficult.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/16-01i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Exhaust all other optimizations before unrolling loops.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Exhaust all other optimizations before unrolling loops.</I></SMALL>
</TABLE>
<H3><A NAME="Heading5"></A><FONT COLOR="#000077">Challenges and Hazards</FONT></H3>
<P>The challenge I put to the readers of <I>PC TECHNIQUES</I> was to write a faster module to replace Listing 16.4. The author of the code that counted the words in my secret test file fastest on my 20 MHz cached 386 would be the winner and receive Numerous Valuable Prizes.</P>

View file

@ -50,14 +50,14 @@
</PRE>
<!-- END CODE SNIP //-->
<P>Harmless enough, save for two things. First, EBX happened to be zero at this point (a leftover from an earlier version of the code, as it turned out), so it was superfluous as a memory-addressing component; this made it possible to use base-only addressing (<B>[EAX]</B>) rather than base&#43;index addressing (<B>[EBX&#43;EAX]</B>), which saves a cycle on the 386. Second: Changing the instruction to <B>CMP [EAX],DH</B> saved 2 cycles&#151;just enough, by good fortune, to speed up the whole program by 5 percent.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/16-02i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I><B>CMP reg,[mem]</B> takes 6 cycles on the 386, but <B>CMP [ mem ],reg</B> takes only 5 cycles; you should always perform<B>CMP</B> with the memory operand on the left on the 386.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I><B>CMP reg,[mem]</B> takes 6 cycles on the 386, but <B>CMP [ mem ],reg</B> takes only 5 cycles; you should always perform<B>CMP</B> with the memory operand on the left on the 386.</I></SMALL>
</TABLE>
<P>(Granted, <B>CMP [<I>mem</I>],<I>reg</I></B> is 1 cycle slower than <B>CMP <I>reg</I>,[<I>mem</I>]</B> on the 286, and they&#146;re both the same on the 8088; in this case, though, the code was specific to the 386. In case you&#146;re curious, both forms take 2 cycles on the 486; quite a lot faster, eh?)</P>
<H4 ALIGN="LEFT"><A NAME="Heading7"></A><FONT COLOR="#000077">Watch Out for Luggable Assumptions!</FONT></H4>
<P>The first lesson to be learned here is not to lug assumptions that may no longer be valid from the 8088/286 world into the wonderful new world of 386 native-mode programming. The second lesson is that after you&#146;ve slaved over your code for a while, you&#146;re in no shape to see its flaws, or to be able to get the new perspectives needed to speed it up. I&#146;ll bet Terje looked at that <B>[EBX&#43;EAX]</B> addressing a hundred times while trying to speed up his code, but he didn&#146;t really see what it did; instead, he saw what it was supposed to do. Mental shortcuts like this are what enable us to deal with the complexities of assembly language without overloading after about 20 instructions, but they can be a major problem when looking over familiar code.</P>
<P>The third, and most interesting, lesson is that a far more fruitful optimization came of all this, one that nicely illustrates that cycle counting is not the key to happiness, riches, and wondrous performance. After getting my 5 percent speedup, I mentioned to Terje the possibility of using a 64K lookup table. (This predated the arrival of entries for the optimization contest.) He said that he had considered it, but it didn&#146;t seem to him to be worthwhile. He couldn&#146;t shake the thought, though, and started to poke around, and one day, <I>voila,</I> he posted a new version of his word count program, WC50, that was <I>much</I> faster than the old version. I don&#146;t have exact numbers, but Terje&#146;s preliminary estimate was 80 percent faster, and word counting&#151;<I>including</I> disk cache access time&#151;proceeds at more than 3 MB per second on a 33 MHz 486. Even allowing for the speed of the 486, those are very impressive numbers indeed.</P>
<P>The point I want to make, though, is that the biggest optimization barrier that Terje faced was that he <I>thought</I> he had the fastest code possible. Once he opened up the possibility that there were faster approaches, and looked beyond the specific approach that he had so carefully optimized, he was able to come up with code that was a <I>lot</I> faster. Consider the incongruity of Terje&#146;s willingness to consider a 5 percent speedup significant in light of his later near-doubling of performance.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/16-03i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Don&#146;t get stuck in the rut of instruction-by-instruction optimization. It&#146;s useful in key loops, but very often, a change in approach will work far greater wonders than any amount of cycle counting can.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Don&#146;t get stuck in the rut of instruction-by-instruction optimization. It&#146;s useful in key loops, but very often, a change in approach will work far greater wonders than any amount of cycle counting can.</I></SMALL>
</TABLE>
<P>By the way, Terje&#146;s WC50 program is a full-fledged counting program; it counts characters, words, and lines, can handle multiple files, and lets you specify the characters that separate words, should you so desire. Source code is provided as part of the archive WC50 comes in. All in all, it&#146;s a nice piece of work, and you might want to take a look at it if you&#146;re interested in really fast assembly code. I wouldn&#146;t call it the <I>fastest</I> word-counting code, though, because I would of course never be so foolish as to call <I>anything</I> the fastest.</P>
<H3><A NAME="Heading8"></A><FONT COLOR="#000077">The Astonishment of Right-Brain Optimization</FONT></H3>

View file

@ -39,7 +39,7 @@
<H3><A NAME="Heading9"></A><FONT COLOR="#000077">Levels of Optimization</FONT></H3>
<P>Three levels of optimization were evident in the word-counting entries I received in response to my challenge. I&#146;d briefly describe them as &#147;fine-tuning,&#148; &#147;new perspective,&#148; and &#147;table-driven state machine.&#148; The latter categories produce faster code, but, by the same token, they are harder to design, harder to implement, and more difficult to understand, so they&#146;re suitable for only the most demanding applications. (Heck, I don&#146;t even guarantee that David Stafford&#146;s entry works perfectly, although, knowing him, it probably does; the more complex and cryptic the code, the greater the chance for obscure bugs.)
</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/16-04i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Remember, optimize only when needed, and stop when further optimization will not be noticed. Optimization that&#146;s not perceptible to the user is like buying Telly Savalas a comb; it&#146;s not going to do any harm, but it&#146;s nonetheless a waste of time.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Remember, optimize only when needed, and stop when further optimization will not be noticed. Optimization that&#146;s not perceptible to the user is like buying Telly Savalas a comb; it&#146;s not going to do any harm, but it&#146;s nonetheless a waste of time.</I></SMALL>
</TABLE>
<H4 ALIGN="LEFT"><A NAME="Heading10"></A><FONT COLOR="#000077">Optimization Level 1: Good Code</FONT></H4>
<P>The first level of optimization involves fine-tuning and clever use of the instruction set. The basic framework is still the same as my code (which in turn is basically the same as that of the original C code), but that framework is implemented more efficiently.

View file

@ -96,12 +96,12 @@
<TD COLSPAN="4"><HR>
</TABLE>
<P>I was wrong. Wrong, wrong, wrong. (But at least I was smart enough to use a profiler before actually writing any new code.) Table 17.1 shows where the time actually goes in Listings 17.1 and 17.2. As you can see, the time taken by <B>draw_pixel(),</B> <B>copy_cells(),</B> and <I>everything</I> other than calculating the next generation is nothing more than noise. We could optimize these routines right down to executing <I>instantaneously,</I> and you know what? It wouldn&#146;t make the slightest perceptible difference in how fast the program runs. Given the present state of our Game of Life implementation, the only areas worth looking at for possible optimizations are <B>cell_state()</B> and <B>next_generation().</B></P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/17-01i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><I><SMALL>It&#146;s worth noting, though, that one reason <B>draw_pixel()</B> doesn&#146;t much affect performance is that in Listing 17.1, we&#146;re smart enough to redraw pixels only when their states change, rather than during every generation. Detecting and eliminating redundant operations is part of knowing the nature of your data, and is a potent optimization technique that will be extremely useful a little later in this chapter.</SMALL></I>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><I><SMALL>It&#146;s worth noting, though, that one reason <B>draw_pixel()</B> doesn&#146;t much affect performance is that in Listing 17.1, we&#146;re smart enough to redraw pixels only when their states change, rather than during every generation. Detecting and eliminating redundant operations is part of knowing the nature of your data, and is a potent optimization technique that will be extremely useful a little later in this chapter.</SMALL></I>
</TABLE>
<H3><A NAME="Heading6"></A><FONT COLOR="#000077">The Hazards and Advantages of Abstraction</FONT></H3>
<P>How can we speed up <B>cell_state()</B> and <B>next_generation()</B>? I&#146;ll tell you how <I>not</I> to do it: By writing those member functions in assembly. It&#146;s tempting to say that <B>cell_state()</B> is taking all the time, so we need to speed it up with assembly, but what we really need to do is figure out <I>why</I> <B>cell_state()</B> is taking all the time, then address that aspect of the program directly.</P>
<P>Once you know where you need to optimize, the one word to keep in mind isn&#146;t assembly, it&#146;s...plastics. No, actually, it&#146;s <I>abstraction.</I> Well-written C and especially C<SMALL>&#43;&#43;</SMALL> programs are highly abstract models. For example, Listing 17.1 essentially creates a new programming language in which cells are tangible things, with built-in manipulation instructions. Given the cellmap member functions, you don&#146;t even need to know the cell storage format! This is a wonderful thing, in general; it saves programming time and bugs, and frees you to work on the application&#146;s needs, rather than implementation details.</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/17-02i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><I><SMALL>However, if you never look beneath the surface of the abstract model at the implementation details, you have no idea of what the true performance cost of various operations</I><I> is, and, without that, you have largely surrendered control over performance.</SMALL></I>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><I><SMALL>However, if you never look beneath the surface of the abstract model at the implementation details, you have no idea of what the true performance cost of various operations</I><I> is, and, without that, you have largely surrendered control over performance.</SMALL></I>
</TABLE>
<P>Having said that, let me hasten to add that algorithmic improvements can make a big difference even when working at a purely abstract level. For a large unordered data set, a high-level Quicksort will beat the pants off the best-implemented insertion sort you can imagine. Still, you can optimize your algorithm from here &#146;til doomsday, and if you have a fast algorithm running on top of a highly abstract programming model, you&#146;ll almost certainly end up with a slow program. In Listing 17.1, the abstraction that&#146;s killing us is that of looking at the eight neighbors with eight completely independent operations, requiring eight calls to <B>cell_state()</B> and eight calculations of cell address and cell mask. In fact, given the nature of cell storage, the eight neighbors are in a fixed relationship to one another, and the addresses and masks of all eight can generally be found very easily via hard-wired offsets and shifts once the address and mask of any one is known.</P><P><BR></P>
<CENTER>

View file

@ -130,7 +130,7 @@ neighbor_count&#43;&#43;;
<P>The neighbor-counting code is brought into <B>next_generation,</B> eliminating many function calls and from-scratch address/mask calculations; all multiplies are eliminated by using pointers and addition; and all cells are accessed directly via pointers and masks, eliminating all remaining function calls and from-scratch address/mask calculations.</P>
<P>The net effect of these optimizations is that Listing 17.4 is more than twice as fast as Listing 17.3; we&#146;ve achieved the desired 18 generations per second, albeit only on a 486, and only at 96&#215;96. (The <B>#define</B> that enables code limiting the speed to 18 Hz, which seemed ridiculous in Listing 17.1, is actually useful for keeping the generations from iterating too quickly when Listing 17.4 is running on a 486, especially with a small cellmap like 48&#215;48.) We&#146;ve sped things up by about eight times so far; we need to increase our speed another ten times to reach our goal of 200&#215;200 at 18 generations per second on a 20 MHz 386.</P>
<P>It&#146;s undoubtedly possible to improve the performance of Listing 17.4 further by fine-tuning the code, but no tremendous improvement is possible that way.</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/17-03i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><I><SMALL>Once you&#146;ve reached the point of fine-tuning pointer usage and register variables and the like in C or C<SMALL>&#43;&#43;</SMALL>, you&#146;ve become compiler-dependent; you therefore might as well go to assembly and get the real McCoy.</SMALL></I>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><I><SMALL>Once you&#146;ve reached the point of fine-tuning pointer usage and register variables and the like in C or C<SMALL>&#43;&#43;</SMALL>, you&#146;ve become compiler-dependent; you therefore might as well go to assembly and get the real McCoy.</SMALL></I>
</TABLE>
<P><BR></P>
<CENTER>

View file

@ -51,7 +51,7 @@
<P>Earlier in this chapter, we looked at a straightforward Game of Life implementation, then increased performance considerably by making the implementation a little less abstract and a little less general. We made a small change to the cellmap format, adding padding bytes off the edges so that pointer arithmetic would always work, but the major optimizations were moving the critical code into a single loop and using pointers rather than member functions whenever possible. In other words, we took what we already knew and made it more efficient.
</P>
<P>Now it&#146;s time to re-examine the nature of this programming task from the ground up, looking for things that we <I>don&#146;t</I> yet know. Let&#146;s take a moment to review what the Game of Life consists of. The basic task is evolving a new generation, and that&#146;s done by looking at the number of &#147;on&#148; neighbors a cell has and the cell&#146;s own state. If a cell is on, and two or three neighbors are on, then the cell stays on; otherwise, an on-cell is turned off. If a cell is off and exactly three neighbors are on, then the cell is turned on; otherwise, an off-cell stays off. That&#146;s all there is to it. As any fool can see, the trick is to arrange things so that we can count neighbors and check the cell state as quickly as possible. Large lookup tables, oddly encoded cellmaps, and lots of bit-twiddling assembly code spring to mind as possible approaches. Can&#146;t you just feel your adrenaline start to pump?</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/17-04i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><I><SMALL>Relax. Step back. Try to divine the true nature of the problem. The object is not to count neighbors and check cell states as quickly as possible; that&#146;s just one possible implementation. The object is to determine when a cell&#146;s state must be changed and to change it appropriately, and that&#146;s what we need to do as quickly as possible.</SMALL></I>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><I><SMALL>Relax. Step back. Try to divine the true nature of the problem. The object is not to count neighbors and check cell states as quickly as possible; that&#146;s just one possible implementation. The object is to determine when a cell&#146;s state must be changed and to change it appropriately, and that&#146;s what we need to do as quickly as possible.</SMALL></I>
</TABLE>
<P>What difference does that new perspective make? Let&#146;s approach it this way. What does a typical cellmap look like? As it happens, after a few generations, the vast majority of cells are off. In fact, the vast majority of cells are not only off but are entirely surrounded by off-cells. Also, cells change state infrequently; in any given generation after the first few, most cells remain in the same state as in the previous generation.
</P>

View file

@ -41,7 +41,7 @@
<P>Anyway, using the large model helps illustrate that it&#146;s the data representation and the data processing approach you choose that matter most. Optimization details like memory models and segments and in-line functions and assembly language are important but secondary. Let your mind roam creatively before you start coding. Otherwise, you may find you&#146;re writing well-tuned slow code, which is by no means the same thing as fast code.</P>
<P>Take a close look at Listing 17.5. You will see that it&#146;s quite a bit simpler than Listing 17.4. To some extent, that&#146;s because I decided to hard-wire the program to wrap around from one edge of the cellmap to the other (it&#146;s much more interesting that way), but the main reason is that it&#146;s a lot easier to work with the neighbor-count model. There&#146;s no complex mask and pointer management, and the only thing that <I>really</I> needs to be optimized is scanning for zero bytes. (And, in fact, I haven&#146;t optimized even that because it&#146;s done in a C<SMALL>&#43;&#43;</SMALL> loop; it should really be <B>REPZ SCASB.</B>)</P>
<P>In truth, none of the code in Listing 17.5 is particularly well-optimized, and, as I noted, the program must be compiled with the large model for large cellmaps. Also, of course, the entire program is still in C<SMALL>&#43;&#43;</SMALL>; note well that there&#146;s not a whit of assembly here.</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/17-05i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><I><SMALL>We&#146;ve gotten more than a 30-times speedup simply by removing a little of the abstraction that C<SMALL>&#43;&#43;</SMALL> encourages, and by storing and processing the data in a manner appropriate for the typical nature of the data itself. In other words, we&#146;ve done some linear, left-brained optimization (using pointers and reducing calls) and some non-linear, right-brained optimization (understanding the real problem and listening for the creative whisper of non-obvious solutions).</SMALL></I>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><I><SMALL>We&#146;ve gotten more than a 30-times speedup simply by removing a little of the abstraction that C<SMALL>&#43;&#43;</SMALL> encourages, and by storing and processing the data in a manner appropriate for the typical nature of the data itself. In other words, we&#146;ve done some linear, left-brained optimization (using pointers and reducing calls) and some non-linear, right-brained optimization (understanding the real problem and listening for the creative whisper of non-obvious solutions).</SMALL></I>
</TABLE>
<P>No doubt we could get another two to five times improvement with good assembly code&#151;but that&#146;s dwarfed by a 30-times improvement, so optimization at a conceptual level <I>must</I> come first.</P>
<H4 ALIGN="LEFT"><A NAME="Heading11"></A><FONT COLOR="#000077">The Challenge That Ate My Life</FONT></H4>

View file

@ -43,7 +43,7 @@
</P>
<P>Actually, it wasn&#146;t. The object was to see who could run the fastest according to the limitations placed upon the contest. This is a crucial distinction, although usually taken for granted. Would it have been legitimate if I had cut across the middle of the field? If I had ridden a bike? If I had broken the world record for the 100 meters by dropping 100 meters from a plane? Competition has meaning only within a carefully circumscribed arena.</P>
<P>Why am I telling you this? First, because it is a useful lesson for programming.</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/18-01i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>All programming is performed within limitations, some of which can be bent or changed, but many of which cannot. You cannot change the maximum memory bandwidth of a VGA, or the maximum instruction execution rate of a 486. That is why the stunning 3D demos you see at SIGGRAPH have only passing relevance to everyday life on the desktop. A rule that Intel&#146;s chip designers cannot break is 8086 compatibility, much as I&#146;m sure they&#146;d like to, but of course the flip side is that although RISC chips are technically superior, they command but a small fraction of the market; raw performance is not the arena of competition. Similarly, you will often be unable to change the specifications for the software you implement.</I></SMALL>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>All programming is performed within limitations, some of which can be bent or changed, but many of which cannot. You cannot change the maximum memory bandwidth of a VGA, or the maximum instruction execution rate of a 486. That is why the stunning 3D demos you see at SIGGRAPH have only passing relevance to everyday life on the desktop. A rule that Intel&#146;s chip designers cannot break is 8086 compatibility, much as I&#146;m sure they&#146;d like to, but of course the flip side is that although RISC chips are technically superior, they command but a small fraction of the market; raw performance is not the arena of competition. Similarly, you will often be unable to change the specifications for the software you implement.</I></SMALL>
</TABLE>
<H3><A NAME="Heading3"></A><FONT COLOR="#000077">Breaking the Rules</FONT></H3>
<P>The other reason for the anecdote has to do with the way my second Optimization Challenge worked itself out. If you&#146;ll recall from the last chapter, the challenge I made to the readers of <I>PC TECHNIQUES</I> was to devise the fastest possible version of the Game of Life cellular automata simulation game. I gave an example, laid out the rules, and stood aside. Good thing, too. <I>Apres moi, le deluge....</I></P>

View file

@ -37,14 +37,14 @@
</CENTER>
<P><BR></P>
<TABLE WIDTH="100%">
<TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/19-01i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>The Pentium, on the other hand, has two separate 8K caches, one for code and one for data, so code prefetches can never collide with data fetches; the prefetch queue can stall only when the code being fetched isn&#146;t in the internal code cache.</I></SMALL>
<TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>The Pentium, on the other hand, has two separate 8K caches, one for code and one for data, so code prefetches can never collide with data fetches; the prefetch queue can stall only when the code being fetched isn&#146;t in the internal code cache.</I></SMALL>
</TABLE>
<P>(And yes, self-modifying code still works; as with all Pentium changes, the dual caches introduce no incompatibilities with 386/486 code.) Also, because the code and data caches are separate, code can&#146;t be driven out of the cache in a tight loop that accesses a lot of data, unlike the 486. In addition, the Pentium expands the 486&#146;s 32-byte prefetch queue to 128 bytes. In conjunction with the branch prediction feature (described next), which allows the Pentium to prefetch properly at most branches, this larger prefetch queue means that the Pentium&#146;s two pipes should be better fed than those of any previous x86 processor.
</P>
<H4 ALIGN="LEFT"><A NAME="Heading5"></A><FONT COLOR="#000077">Crossing Cache Lines</FONT></H4>
<P>There are three other characteristics of the Pentium that make for a healthy supply of instruction bytes. One is that the Pentium can prefetch instructions across cache lines. Unlike the 486, where there is a 3-cycle penalty for branching to an instruction that spans a cache line, there&#146;s no such penalty on the Pentium. The second is that the cache line size (the number of bytes fetched from the external cache or main memory on a cache miss) on the Pentium is 32 bytes, twice the size of the 486&#146;s cache line, so a cache miss causes a longer run of instructions to be placed in the cache than on the 486. The third is that the Pentium&#146;s external bus is twice as wide as the 486&#146;s, at 64 bits, and runs twice as fast, at 66 MHz, so the Pentium can fetch both instruction and data bytes from the external cache four times as fast as the 486.
</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/19-02i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>Even when the Pentium is running flat-out with both pipes in use, it can generally consume only about twice as many bytes as the 486; so the ratio of external memory bandwidth to processing power is much improved, although real-world performance is heavily dependent on the size and speed of the external cache.</I></SMALL>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>Even when the Pentium is running flat-out with both pipes in use, it can generally consume only about twice as many bytes as the 486; so the ratio of external memory bandwidth to processing power is much improved, although real-world performance is heavily dependent on the size and speed of the external cache.</I></SMALL>
</TABLE>
<P>The upshot of all this is that at the same clock speed, with code and data that are mostly in the internal caches, the Pentium maxes out somewhere around twice as fast as a 486. (When the caches are missed a lot, the Pentium can get as much as three to four times faster, due to the superior external bus and bigger caches.) Most of this won&#146;t affect how you program, but it is useful to know that you don&#146;t have to worry about instruction fetching. It&#146;s also useful to know the sizes of the caches because a high cache hit rate is crucial to Pentium performance. Cache misses are vastly slower than cache hits (anywhere from two to 50 or more times as slow, depending on the speed of the external cache and whether the external cache misses as well), and the Pentium can&#146;t use the V-pipe on code that hasn&#146;t already been executed out of the cache at least once. This means that it is <I>very</I> important to get the working sets of critical loops to fit in the internal caches.</P>
<P>One change in the Pentium that you definitely do have to worry about is superscalar execution. Utilization of the V-pipe can range from near zero percent to 100 percent, depending on the code being executed, and careful rearrangement of code can have amazing effects. Maxing out V-pipe use is not a trivial task; I&#146;ll spend all of the next chapter discussing it so as to have time to cover it properly. In the meantime, two good references for superscalar programming and other Pentium information are Intel&#146;s <I>Pentium Processor User&#146;s Manual: Volume 3: Architecture and Programming Manual</I> (ISBN 1-55512-195-0; Intel order number 241430-001), and the article &#147;Optimizing Pentium Code&#148; by Mike Schmidt, in <I>Dr. Dobb&#146;s Journal</I> for January 1994.</P>

View file

@ -48,11 +48,11 @@ mov eax,[ebx]
<!-- END CODE SNIP //-->
<P>costs three cycles. On the other hand, as noted above, branch targets can now span cache lines with impunity, so on the Pentium there&#146;s no good argument for the paragraph (that is, 16-byte) alignment that Intel recommends for 486 jump targets. The 32-byte alignment might make for slightly more efficient Pentium cache usage, but would make code much bigger overall.
</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/19-03i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>In fact, given that most jump targets aren&#146;t in performance-critical code, it&#146;s hard to make a compelling argument for aligning branch targets even on the 486. I&#146;d say that no alignment (except possibly where you know a branch target lies in a key loop), or at most dword alignment (for the 386) is plenty, and can shrink code size considerably.</I></SMALL>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>In fact, given that most jump targets aren&#146;t in performance-critical code, it&#146;s hard to make a compelling argument for aligning branch targets even on the 486. I&#146;d say that no alignment (except possibly where you know a branch target lies in a key loop), or at most dword alignment (for the 386) is plenty, and can shrink code size considerably.</I></SMALL>
</TABLE>
<P>Instruction prefixes are awfully expensive; avoid them if you can. (These include size and addressing prefixes, segment overrides, <B>LOCK</B>, and the 0FH prefixes that extend the instruction set with instructions such as <B>MOVSX</B>. The exceptions are conditional jumps, a fast special case.) At a minimum, a prefix byte generally takes an extra cycle and shuts down the V-pipe for that cycle, effectively costing as much as two normal instructions (although prefix cycles can overlap with previous multicycle instructions, or AGIs, as on the 486). This means that using 32-bit addressing or 32-bit operands in a 16-bit segment, or vice versa, makes for bigger code that&#146;s significantly slower. So, for example, you should generally avoid 16-bit variables (shorts, in C) in 32-bit code, although if using 32-bit variables where they&#146;re not needed makes your data space get a lot bigger, you may want to stick with shorts, especially since longs use the cache less efficiently than shorts. The trade-off depends on the amount of data and the number of instructions that reference that data. (eight-bit variables, such as chars, have no extra overhead and can be used freely, although they may be less desirable than longs for compilers that tend to promote variables to longs when performing calculations.) Likewise, you should if possible avoid putting data in the code segment and referring to it with a CS: prefix, or otherwise using segment overrides.</P>
<P><B>LOCK</B> is a particularly costly instruction, especially on multiprocessor machines, because it locks the bus and requires that the hardware be brought into a synchronized state. The cost varies depending on the processor and system, but <B>LOCK</B> can make an <B>INC [<I>mem</I>]</B> instruction (which normally takes 3 cycles) 5, 10, or more cycles slower. Most programmers will never use <B>LOCK</B> on purpose&#151;it&#146;s primarily an operating system instruction&#151;but there&#146;s a hidden gotcha here because the <B>XCHG</B> instruction always locks the bus when used with a memory operand.</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/19-04i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I><B>XCHG</B> is a tempting instruction that&#146;s often used in assembly language; for example, exchanging with video memory is a popular way to read and write VGA memory in a single instruction&#151;but it&#146;s now a bad idea. As it happens, on the 486 and Pentium, using <B>MOV</B>s to read and write memory is faster, anyway; and even on the 486, my measurements indicate a five-cycle tax for <B>LOCK</B> in general, and a nine-cycle execution time for <B>XCHG</B> with memory. Avoid <B>XCHG</B> with memory if you possibly can.</I></SMALL>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I><B>XCHG</B> is a tempting instruction that&#146;s often used in assembly language; for example, exchanging with video memory is a popular way to read and write VGA memory in a single instruction&#151;but it&#146;s now a bad idea. As it happens, on the 486 and Pentium, using <B>MOV</B>s to read and write memory is faster, anyway; and even on the 486, my measurements indicate a five-cycle tax for <B>LOCK</B> in general, and a nine-cycle execution time for <B>XCHG</B> with memory. Avoid <B>XCHG</B> with memory if you possibly can.</I></SMALL>
</TABLE>
<P>As with the 486, don&#146;t use <B>ENTER</B> or <B>LEAVE</B>, which are slower than the equivalent discrete instructions. Also, start using <B>TEST <I>reg,reg</I></B> instead of <B>AND <I>reg,reg</I></B> or <B>OR <I>reg,reg</I></B> to test whether a register is zero. The reason, as we&#146;ll see in Chapter 21, is that <B>TEST</B>, unlike <B>AND</B> and <B>OR</B>, never modifies the target register. Although in this particular case <B>AND</B> and <B>OR</B> don&#146;t modify the target register either, the Pentium has no way of knowing that ahead of time, so if <B>AND</B> or <B>OR</B> goes through the U-pipe, the Pentium may have to shut down the V-pipe for a cycle to avoid potential dependencies on the result of the <B>AND</B> or <B>OR</B>. <B>TEST</B> suffers from no such potential dependencies.</P><P><BR></P>
<CENTER>

View file

@ -38,13 +38,13 @@
<P><BR></P>
<H3><A NAME="Heading8"></A><FONT COLOR="#000077">Branch Prediction</FONT></H3>
<P>One brand-spanking-new feature of the Pentium is <I>branch prediction</I>, whereby the Pentium tries to guess, based on past history, which way (or, for conditional jumps, whether or not), your code will jump at each branch, and prefetches along the likelier path. If the guess is correct, the branch or fall-through takes only 1 cycle&#151;2 cycles less than a branch and the same as a fall-through on the 486; if the guess is wrong, the branch or fall-through takes 4 or 5 cycles (if it executes in the U- or V-pipe, respectively)&#151;1 or 2 cycles more than a branch and 3 or 4 cycles more than a fall-through on the 486.</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/19-05i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>Branch prediction is unprecedented in the x86, and fundamentally alters the nature of pedal-to-the-metal optimization, for the simple reason that it renders unrolled loops largely obsolete. Rare indeed is the loop that can&#146;t afford to spare even 1 or 0 (yes, zero!) cycles per iteration for loop counting, and that&#146;s how low the cost can go for maintaining a loop on the Pentium.</I></SMALL>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>Branch prediction is unprecedented in the x86, and fundamentally alters the nature of pedal-to-the-metal optimization, for the simple reason that it renders unrolled loops largely obsolete. Rare indeed is the loop that can&#146;t afford to spare even 1 or 0 (yes, zero!) cycles per iteration for loop counting, and that&#146;s how low the cost can go for maintaining a loop on the Pentium.</I></SMALL>
</TABLE>
<P>Also, unrolled loops are bigger than normal loops, so there are extra (and expensive) cache misses the first time through the loop if the entire loop isn&#146;t already in the cache; then, too, an unrolled loop will shoulder other code out of the internal and external caches. If in a critical loop you absolutely need the time taken by the loop control instructions, or if you need an extra register that can be freed by unrolling a loop, then by all means unroll the loop. Don&#146;t expect the sort of speed-up you get from this on the 486 or especially the 386, though, and watch out for the cache effects.
</P>
<P>You may well wonder exactly <I>when</I> the Pentium correctly predicts branching. Alas, this is one area that Intel has declined to document, beyond saying that you should endeavor to fall through branches when you have a choice. That&#146;s good advice on every other x86 processor, anyway, so it&#146;s well worth following. Also, it&#146;s a pretty safe bet that in a tight loop, the Pentium will start guessing the right branch direction at the bottom of the loop pretty quickly, so you can treat loop branches as one-cycle instructions.</P>
<P>It&#146;s an equally safe bet that it&#146;s a bad move to have in a loop a conditional branch that goes both ways on a random basis; it&#146;s hard to see how the Pentium could consistently predict such branches correctly, and mispredicted branches are more expensive than they might appear to be. Not only does a mispredicted branch take 4 or 5 cycles, but the Pentium can potentially execute as many as 8 or 10 instructions in that time&#151;3 times as many as the 486 can execute during its branch time&#151;so correct branch prediction (or eliminating branch instructions, if possible) is very important in inner loops. Note that on the 486 you can count on a branch to take 1 cycle when it falls through, but on the Pentium you can&#146;t be sure whether it will take 1 or either 4 or 5 cycles on any given iteration.</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/19-06i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>As things currently stand, branch prediction is an annoyance for assembly language optimization because it&#146;s impossible to be certain exactly how code will perform until you measure it, and even then it&#146;s difficult to be sure exactly where the cycles went. All I can say is try to fall through branches if possible, and try to be consistent in your branching if not.</I></SMALL>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>As things currently stand, branch prediction is an annoyance for assembly language optimization because it&#146;s impossible to be certain exactly how code will perform until you measure it, and even then it&#146;s difficult to be sure exactly where the cycles went. All I can say is try to fall through branches if possible, and try to be consistent in your branching if not.</I></SMALL>
</TABLE>
<H3><A NAME="Heading9"></A><FONT COLOR="#000077">Miscellaneous Pentium Topics</FONT></H3>
<P>The Pentium has all the instructions of the 486, plus a few new ones. One much-needed instruction that has finally made it into the instruction set is <B>CPUID</B>, which allows your code to determine what processor it&#146;s running on. <B>CPUID</B> is 15 years late, but at least it&#146;s finally here. Another new instruction is <B>CMPXCHG8B</B>, which does a compare and conditional exchange on a qword. <B>CMPXCHG8B</B> doesn&#146;t seem to me to be a particularly useful instruction, but I&#146;m sure Intel wouldn&#146;t have added it without a reason; if you know of a use for it, please pass it along to me.</P>

View file

@ -53,7 +53,7 @@
<BR><A HREF="javascript:displayWindow('images/20-01.jpg',407,335)"> --><FONT COLOR="#000077"><B>Figure 20.1</B></FONT></A>&nbsp;&nbsp;<I>The Pentium&#146;s two pipes.</I>
</P>
<P>Getting two instructions executing simultaneously in the two pipes is trickier than it sounds, not only because the V-pipe can handle only a relatively small subset of the Pentium&#146;s instruction set, but also because those instructions that the V-pipe can handle are able to pair only with certain U-pipe instructions. For example, <B>MOVSD</B> uses both pipes, so no instruction can be executed in parallel with <B>MOVSD</B>.</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/20-01i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>The use of both pipes does make <B>MOVSD</B> nearly twice as fast on the Pentium as on the 486, but it&#146;s nonetheless slower than using equivalent simpler instructions that allow for superscalar execution. Stick to the Pentium&#146;s RISC-like instructions&#151;the pairable instructions I&#146;ll discuss next&#151;when you&#146;re seeking maximum performance, with just a few exceptions such as <B>REP MOVS</B> and <B>REP STOS</B>.</I></SMALL>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>The use of both pipes does make <B>MOVSD</B> nearly twice as fast on the Pentium as on the 486, but it&#146;s nonetheless slower than using equivalent simpler instructions that allow for superscalar execution. Stick to the Pentium&#146;s RISC-like instructions&#151;the pairable instructions I&#146;ll discuss next&#151;when you&#146;re seeking maximum performance, with just a few exceptions such as <B>REP MOVS</B> and <B>REP STOS</B>.</I></SMALL>
</TABLE>
<P>Trickier yet, register contention can shut down the V-pipe on any given cycle, and Address Generation Interlocks (AGIs) can stall either pipe at any time, as we&#146;ll see in the next chapter.
</P>

View file

@ -123,7 +123,7 @@ ROL/ROR/RCL/RCR reg,1 (1 cycle)
</PRE>
<!-- END CODE //-->
<P><B>Table 20.2 Instructions that, when executed in the U-pipe, allow V-pipe-executable instructions to execute simultaneously (pair) in the V-pipe.</B><HR></P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/20-02i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>A fundamental rule of Pentium optimization is that it pays to break complex instructions into equivalent simple instructions, then shuffle the simple instructions for maximum use of the V-pipe. This is true partly because most of the pairable instructions are simple instructions, and partly because breaking instructions into pieces allows more freedom to rearrange code to avoid the AGIs and register contention I&#146;ll discuss in the next chapter.</I></SMALL>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>A fundamental rule of Pentium optimization is that it pays to break complex instructions into equivalent simple instructions, then shuffle the simple instructions for maximum use of the V-pipe. This is true partly because most of the pairable instructions are simple instructions, and partly because breaking instructions into pieces allows more freedom to rearrange code to avoid the AGIs and register contention I&#146;ll discuss in the next chapter.</I></SMALL>
</TABLE>
<P><A NAME="Fig2"><!-- </A><A HREF="javascript:displayWindow('images/20-02.jpg',405,144 )"> --><IMG SRC="images/20-02.jpg"><BR><!-- </A>
<BR><A HREF="javascript:displayWindow('images/20-02.jpg',405,144)"> --><FONT COLOR="#000077"><B>Figure 20.2</B></FONT></A>&nbsp;&nbsp;<I>Instruction flow through the two pipes.</I>
@ -163,7 +163,7 @@ mov [MemVar],edx
</PRE>
<!-- END CODE SNIP //-->
<P>The single complex instruction takes 3 cycles and is 6 bytes long; with proper sequencing, interleaving the simple instructions with other instructions that don&#146;t use EDX or <B>Mem Var</B>, the three-instruction sequence can be reduced to 1.5 cycles, but it is <I>14</I> bytes long.</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/20-03i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>It&#146;s not unusual for Pentium optimization to approximately double both performance and code size at the same time. In an important loop, go for performance and ignore the size, but on a program-wide basis, the size bears watching.</I></SMALL>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>It&#146;s not unusual for Pentium optimization to approximately double both performance and code size at the same time. In an important loop, go for performance and ignore the size, but on a program-wide basis, the size bears watching.</I></SMALL>
</TABLE>
<P><BR></P>
<CENTER>

View file

@ -39,7 +39,7 @@
<H3><A NAME="Heading5"></A><FONT COLOR="#000077">Lockstep Execution</FONT></H3>
<P>You may wonder why anyone would bother breaking <B>ADD [MemVar],EAX</B> into three instructions, given that this instruction can go through either pipe with equal ease. The answer is that while the memory-accessing instructions other than <B>MOV, PUSH</B>, and <B>POP</B> listed in Table 20.1 (that is, <B>INC/DEC [<I>mem</I>], ADD/SUB/XOR/AND/OR/CMP/ADC/SBB <I>reg</I>,[<I>mem</I>]</B>, and <B>ADD/SUB/XOR/AND/OR/CMP/ADC/SBB [<I>mem</I>],<I>reg/immed</I></B>) can be paired, they do not provide the 100 percent overlap that we seek. If you look at Tables 20.1 and 20.2, you will see that instructions taking from 1 to 3 cycles can pair. However, any pair of instructions goes through the two pipes in lockstep. This means, for example, that if <B>ADD [EBX],EDX</B> is going through the U-pipe, and <B>INC EAX</B> is going through the V-pipe, the V-pipe will be idle for 2 of the 3 cycles that the U-pipe takes to execute its instruction, as shown in Figure 20.4. Out of the theoretical 6 cycles of work that can be done during this time, we actually get only 4 cycles of work, or 67 percent utilization. Even though these instructions pair, then, this sequence fails to make maximum use of the Pentium&#146;s horsepower.</P>
<P>The key here is that when two instructions pair, both execution units are tied up until both instructions have finished (which means at least for the amount of time required for the longer of the two to execute, plus possibly some extra cycles for pairable instructions that can&#146;t fully overlap, as described below). The logical conclusion would seem to be that we should strive to pair instructions of the same lengths, but that is often not correct.</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/20-04i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>The actual rule is that we should strive to pair one-cycle instructions (or, at most, two-cycle instructions, but not three-cycle instructions), which in turn leads to the corollary that we should, in general, use mostly one-cycle instructions when optimizing.</I></SMALL>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>The actual rule is that we should strive to pair one-cycle instructions (or, at most, two-cycle instructions, but not three-cycle instructions), which in turn leads to the corollary that we should, in general, use mostly one-cycle instructions when optimizing.</I></SMALL>
</TABLE>
<P><A NAME="Fig4"><!-- </A><A HREF="javascript:displayWindow('images/20-04.jpg',411,224 )"> --><IMG SRC="images/20-04.jpg"><BR><!-- </A>
<BR><A HREF="javascript:displayWindow('images/20-04.jpg',411,224)"> --><FONT COLOR="#000077"><B>Figure 20.4</B></FONT></A>&nbsp;&nbsp;<I>Lockstep execution and idle time in the V-pipe.</I>

View file

@ -67,7 +67,7 @@ mov [ebx],dl
<H4 ALIGN="LEFT"><A NAME="Heading7"></A><FONT COLOR="#000077">Register Starvation</FONT></H4>
<P>The above examples should make it pretty clear that effective superscalar programming puts a lot of strain on the Pentium&#146;s relatively small register set. There are only seven general-purpose registers (I strongly suggest using EBP in critical loops), and it does not help to have to sacrifice one of those registers for temporary storage on each complex memory operation; in pre-superscalar days, we used to employ those handy CISC memory instructions to do all that stuff without using any extra registers.
</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/20-05i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>More problematic still is that for maximum pairing, you&#146;ll typically have two operations proceeding at once, one in each pipe, and trying to keep two operations in registers at once is difficult indeed. There&#146;s not much to be done about this, other than clever and Spartan register usage, but be aware that it&#146;s a major element of Pentium performance programming.</I></SMALL>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>More problematic still is that for maximum pairing, you&#146;ll typically have two operations proceeding at once, one in each pipe, and trying to keep two operations in registers at once is difficult indeed. There&#146;s not much to be done about this, other than clever and Spartan register usage, but be aware that it&#146;s a major element of Pentium performance programming.</I></SMALL>
</TABLE>
<P>Also be aware that prefixes of every sort, with the sole exception of the 0FH prefix on non-short conditional jumps, always execute in the U-pipe, and that Intel&#146;s documentation indicates that no pairing can happen while a prefix byte executes. (As I&#146;ll discuss in the next chapter, my experiments indicate that this rule doesn&#146;t always apply to multiple-cycle instructions, but you still won&#146;t go far wrong by assuming that the above rule is correct and trying to eliminate prefix bytes.) A prefix byte takes one cycle to execute; after that cycle, the actual prefixed instruction itself will go through the U-pipe, and if it and the following instruction are mutually pairable, then they will pair. Nonetheless, prefix bytes are very expensive, effectively taking at least as long as two normal instructions, and possibly, if a prefixed instruction could otherwise have paired in the V-pipe with the previous instruction, taking as long as three normal instructions, as shown in Figure 20.8.
</P>

View file

@ -106,7 +106,7 @@
</P>
<P>The method used in the VGA BIOS to set registers is to point DX to the desired Index register, load AL with the index, perform a byte <B>OUT</B>, increment DX to point to the Data register (except in the case of the AC, where DX remains the same), load AL with the desired data, and perform a byte <B>OUT</B>. A handy shortcut is to point DX to the desired Index register, load AL with the index, load AH with the data, and perform a word <B>OUT</B>. Since the high byte of the <B>OUT</B> value goes to port DX&#43;1, this is equivalent to the first method but is faster. However, this technique does not work for programming the AC Index and Data registers; both AC registers are addressed at 3C0H, so two separate byte <B>OUT</B>s must be used to program the AC. (Actually, word <B>OUT</B>s to the AC do work in the EGA, but not in the VGA, so they shouldn&#146;t be used.) As mentioned above, you must be sure which mode&#151;Index or Data&#151;the AC is in before you do an <B>OUT</B> to 3C0H; you can read the Input Status 1 register at any time to force the AC to Index mode.</P>
<P>How safe is the word-<B>OUT</B> method of addressing VGA registers? I have, in the past, run into adapter/computer combinations that had trouble with word <B>OUT</B>s; however, all such problems I am aware of have been fixed. Moreover, a great deal of graphics software now uses word <B>OUT</B>s, so any computer or VGA that doesn&#146;t properly support word <B>OUT</B>s could scarcely be considered a clone at all.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/23-01i.jpg"><TD WIDTH="95%" VALIGN="TOP"><SMALL><I>A speed tip: The setting of each chip&#146;s Index register remains the same until it is reprogrammed. This means that in cases where you are setting the same internal register repeatedly, you can set the Index register to point to that internal register once, then write to the Data register multiple times. For example, the Bit Mask register (GC register 8) is often set repeatedly inside a loop when drawing lines. The standard code for this is:</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%" VALIGN="TOP"><SMALL><I>A speed tip: The setting of each chip&#146;s Index register remains the same until it is reprogrammed. This means that in cases where you are setting the same internal register repeatedly, you can set the Index register to point to that internal register once, then write to the Data register multiple times. For example, the Bit Mask register (GC register 8) is often set repeatedly inside a loop when drawing lines. The standard code for this is:</I></SMALL>
</TABLE>
<!-- CODE SNIP //-->
<PRE>

View file

@ -49,7 +49,7 @@
<P>Smooth horizontal panning is provided by the Horizontal Pel Panning register, AC register 13H, working in conjunction with the start address. Up to 7 pixels worth of single pixel panning of the displayed image to the left is performed by increasing the Horizontal Pel Panning register from 0 to 7. This exhausts the range of motion possible via the Horizontal Pel Panning register; the next pixel&#146;s worth of smooth panning is accomplished by incrementing the start address by one and resetting the Horizontal Pel Panning register to 0. Smooth horizontal panning should be viewed as a series of fine adjustments in the 8-pixel range between coarse byte-sized adjustments.</P>
<P>A horizontal panning oddity: Alone among VGA modes, text mode (in most cases) has 9 dots per character clock. Smooth panning in this mode requires cycling the Horizontal Pel Panning register through the values 8, 0, 1, 2, 3, 4, 5, 6, and 7. 8 is the &#147;no panning&#148; setting.</P>
<P>There is one annoying quirk about programming the AC. When the AC Index register is set, only the lower five bits are used as the internal index. The next most significant bit, bit 5, controls the source of the video data sent to the monitor by the VGA. When bit 5 is set to 1, the output of the palette RAM, derived from display memory, controls the displayed pixels; this is normal operation. When bit 5 is 0, video data does not come from the palette RAM, and the screen becomes a solid color. The only time bit 5 of the AC Index register should be 0 is during the setting of a palette RAM register, since the CPU is only able to write to palette RAM when bit 5 is 0. (Some VGAs do not enforce this, but you should always set bit 5 to 0 before writing to the palette RAM just to be safe.) Immediately after setting palette RAM, however, 20h (or any other value with bit 5 set to 1) should be written to the AC Index register to restore normal video, and at all other times bit 5 should be set to 1.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/23-02i.jpg"><TD WIDTH="95%" VALIGN="TOP"><SMALL><I>By the way, palette RAM can be set via the BIOS video interrupt (interrupt 10H), function 10H. Whenever an VGA function can be performed reasonably well through a BIOS function, as it can in the case of setting palette RAM, it should be, both because there is no point in reinventing the wheel and because the BIOS may well mask incompatibilities between the IBM VGA and VGA clones.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%" VALIGN="TOP"><SMALL><I>By the way, palette RAM can be set via the BIOS video interrupt (interrupt 10H), function 10H. Whenever an VGA function can be performed reasonably well through a BIOS function, as it can in the case of setting palette RAM, it should be, both because there is no point in reinventing the wheel and because the BIOS may well mask incompatibilities between the IBM VGA and VGA clones.</I></SMALL>
</TABLE>
<P><BR></P>
<CENTER>

View file

@ -38,7 +38,7 @@
<P><BR></P>
<P>One possible solution to this problem is to pick a second page start address that has a 0 value for the lower byte, so only the Start Address High register ever needs to be set, but in the sample program in Listing 23.1 I&#146;ve gone for generality and always set both bytes. To avoid mismatched start address bytes, the sample program waits for pixel data to be displayed, as indicated by the Display Enable status; this tells us we&#146;re somewhere in the displayed portion of the frame, far enough away from vertical sync so we can be sure the new start address will get used at the next vertical sync. Once the Display Enable status is observed, the program sets the new start address, waits for vertical sync to happen, sets the new pel panning state, and then continues drawing. Don&#146;t worry about the details right now; page flipping will come up again, at considerably greater length, in later chapters.
</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/23-03i.jpg"><TD WIDTH="95%" VALIGN="TOP"><SMALL><I>As an interesting side note, be aware that if you run DOS software under a multitasking environment such as Windows NT, timeslicing delays can make mismatched start address bytes or mismatched start address and pel panning settings much more likely, for the graphics code can be interrupted at any time. This is also possible, although much less likely, under non-multitasking environments such as DOS, because strategically placed interrupts can cause the same sorts of problems there. For maximum safety, you should disable interrupts around the key portions of your page-flipping code, although here we run into the problem that if interrupts are disabled from the time we start looking for Display Enable until we set the Pel Panning register, they will be off for far too long, and keyboard, mouse, and network events will potentially be lost. Also, disabling interrupts won&#146;t help in true multitasking environments, which never let a program hog the entire CPU. This is one reason that pel panning, although indubitably flashy, isn&#146;t widely used and should be reserved for only those cases where it&#146;s absolutely necessary.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%" VALIGN="TOP"><SMALL><I>As an interesting side note, be aware that if you run DOS software under a multitasking environment such as Windows NT, timeslicing delays can make mismatched start address bytes or mismatched start address and pel panning settings much more likely, for the graphics code can be interrupted at any time. This is also possible, although much less likely, under non-multitasking environments such as DOS, because strategically placed interrupts can cause the same sorts of problems there. For maximum safety, you should disable interrupts around the key portions of your page-flipping code, although here we run into the problem that if interrupts are disabled from the time we start looking for Display Enable until we set the Pel Panning register, they will be off for far too long, and keyboard, mouse, and network events will potentially be lost. Also, disabling interrupts won&#146;t help in true multitasking environments, which never let a program hog the entire CPU. This is one reason that pel panning, although indubitably flashy, isn&#146;t widely used and should be reserved for only those cases where it&#146;s absolutely necessary.</I></SMALL>
</TABLE>
<P>Waiting for the sync pulse has the side effect of causing program execution to synchronize to the VGA&#146;s frame rate of 60 or 70 frames per second, depending on the display mode. This synchronization has the useful consequence of causing the program to execute at the same speed on any CPU that can draw fast enough to complete the drawing in a single frame; the program just idles for the rest of each frame that it finishes before the VGA is finished displaying the previous frame.
</P>

View file

@ -60,7 +60,7 @@ MOV DX,(VALUE2 SHL 8) OR VALUE1
</PRE>
<!-- END CODE SNIP //-->
<P>which takes only three bytes and is faster, being a single instruction. (Note, though, that in 32-bit protected mode, there&#146;s a size and performance penalty for 16-bit instructions such as the <B>MOV</B> above; see the first part of this book for details.) As shown, a macro is an ideal place to use this technique; the macro invocation can refer to two separate byte values, making matters easier for the programmer, while the macro itself can combine the values into a single word-sized constant.</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/24-01i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>A minor optimization tip illustrated in the listing is the use of <B>INC AX</B> and <B>DEC AX</B> in the <B>DrawVerticalBox</B> subroutine when only AL actually needs to be modified. Word-sized register increment and decrement instructions (or dword-sized instructions in 32-bit protected mode) are only one byte long, while byte-size register increment and decrement instructions are two bytes long. Consequently, when size counts, it is worth using a whole 16-bit (or 32-bit) register instead of the low 8 bits of that register for <B>INC</B> and <B>DEC</B>&#151;if you don&#146;t need the upper portion of the register for any other purpose, or if you can be sure that the <B>INC</B> or <B>DEC</B> won&#146;t affect the upper part of the register.</I></SMALL>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>A minor optimization tip illustrated in the listing is the use of <B>INC AX</B> and <B>DEC AX</B> in the <B>DrawVerticalBox</B> subroutine when only AL actually needs to be modified. Word-sized register increment and decrement instructions (or dword-sized instructions in 32-bit protected mode) are only one byte long, while byte-size register increment and decrement instructions are two bytes long. Consequently, when size counts, it is worth using a whole 16-bit (or 32-bit) register instead of the low 8 bits of that register for <B>INC</B> and <B>DEC</B>&#151;if you don&#146;t need the upper portion of the register for any other purpose, or if you can be sure that the <B>INC</B> or <B>DEC</B> won&#146;t affect the upper part of the register.</I></SMALL>
</TABLE>
<P>The latches and ALUs are central to high-performance VGA code, since they allow programs to process across all four memory planes without a series of <B>OUT</B>s and read/write operations. It is not always easy to arrange a program to exploit this power, however, because the ALUs are far more limited than a CPU. In many instances, however, additional hardware in the VGA, including the bit mask, the set/reset features, and the barrel shifter, can assist the ALUs in controlling data, as we&#146;ll see in the next few chapters.</P><P><BR></P>
<CENTER>

View file

@ -37,7 +37,7 @@
</CENTER>
<P><BR></P>
<P>It&#146;s worth pointing out again that the bit mask operates on the data in the latches, not on the data in display memory. This makes the bit mask a flexible resource that with a little imagination can be used for some interesting purposes. For example, you could fill the latches with a solid background color (by writing the color somewhere in display memory, then reading that location to load the latches), and then use the Bit Mask register (or write mode 3, as we&#146;ll see later) as a mask through which to draw a foreground color stencilled into the background color <I>without</I> reading display memory first. This only works for writing whole bytes at a time (clipped bytes require the use of the bit mask; unfortunately, we&#146;re already using it for stencilling in this case), but it completely eliminates reading display memory and does foreground-plus-background drawing in one blurry-fast pass.</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" WIDTH="5%" ALIGN="LEFT"><IMG SRC="images/25-01i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>This last-described example is a good illustration of how I&#146;d suggest you approach the VGA: As a rich collection of hardware resources that can profitably be combined in some non-obvious ways. Don&#146;t let yourself be limited by the obvious applications for the latches, bit mask, write modes, read modes, map mask, ALUs, and set/reset circuitry. Instead, try to imagine how they could work together to perform whatever task you happen to need done at any given time. I&#146;ve made my code as much as four times faster by doing this, as the discussion of Mode X in Chapters 47&#150;49 demonstrates.</I></SMALL>
<TABLE WIDTH="100%"><TD VALIGN="TOP" WIDTH="5%" ALIGN="LEFT"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>This last-described example is a good illustration of how I&#146;d suggest you approach the VGA: As a rich collection of hardware resources that can profitably be combined in some non-obvious ways. Don&#146;t let yourself be limited by the obvious applications for the latches, bit mask, write modes, read modes, map mask, ALUs, and set/reset circuitry. Instead, try to imagine how they could work together to perform whatever task you happen to need done at any given time. I&#146;ve made my code as much as four times faster by doing this, as the discussion of Mode X in Chapters 47&#150;49 demonstrates.</I></SMALL>
</TABLE>
<P>The example code in Listing 25.1 is designed to illustrate the use of the Data Rotate and Bit Mask registers, and is not as fast or as complete as it might be. The case where text <I>is</I> byte-aligned could be detected and performed much faster, without the use of the Bit Mask or Data Rotate registers and with only one display memory access per font byte (to write the font byte), rather than four (to read display memory and write the font byte to each of the two bytes the character spans). Likewise, non-aligned text drawing could be streamlined to one display memory access per byte by having the CPU rotate and combine the font data directly, rather than setting up the VGA&#146;s hardware to do it. (Listing 25.1 was designed to illustrate VGA data rotation and bit masking rather than the fastest way to draw text. We&#146;ll see faster text-drawing code soon.) One excellent rule of thumb is to minimize display memory accesses of all types, especially reads, which tend to be considerably slower than writes. Also, in Listing 25.1 it would be faster to use a table lookup to calculate the bit masks for the two halves of each character rather than the shifts used in the example.</P>
<P>For another (and more complex) example of drawing bit-mapped text on the VGA, see John Cockerham&#146;s article, &#147;Pixel Alignment of EGA Fonts,&#148; <I>PC Tech Journal</I>, January, 1987. Parenthetically, I&#146;d like to pass along John&#146;s comment about the VGA: &#147;When programming the VGA, <I>everything</I> is complex.&#148;</P>

View file

@ -41,7 +41,7 @@
<H3><A NAME="Heading8"></A><FONT COLOR="#000077">Notes on Set/Reset</FONT></H3>
<P>The set/reset circuitry is not active in write modes 1 or 2. The Enable Set/Reset register is inactive in write mode 3, but the Set/Reset register provides the primary drawing color in write mode 3, as discussed in the next chapter.
</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/25-02i.jpg"><TD WIDTH="95%"><SMALL><I>Be aware that because set/reset directly replaces CPU data, it does not necessarily have to force an entire display memory byte to 0 or 0FFH, even when set/reset is replacing CPU data for all planes. For example, if the Bit Mask register is set to 80H, the set/reset circuitry can only modify bit 7 of the destination byte in each plane, since the other seven bits will come from the latches for each plane. Similarly, the set/reset value for each plane can be modified by that plane&#146;s ALU. Once again, this illustrates that set/reset merely replaces the CPU data for selected planes; the set/reset value is then processed in exactly the same way that CPU data normally is.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>Be aware that because set/reset directly replaces CPU data, it does not necessarily have to force an entire display memory byte to 0 or 0FFH, even when set/reset is replacing CPU data for all planes. For example, if the Bit Mask register is set to 80H, the set/reset circuitry can only modify bit 7 of the destination byte in each plane, since the other seven bits will come from the latches for each plane. Similarly, the set/reset value for each plane can be modified by that plane&#146;s ALU. Once again, this illustrates that set/reset merely replaces the CPU data for selected planes; the set/reset value is then processed in exactly the same way that CPU data normally is.</I></SMALL>
</TABLE>
<H3><A NAME="Heading9"></A><FONT COLOR="#000077">A Brief Note on Word OUTs</FONT></H3>
<P>In the early days of the EGA and VGA, there was considerable debate about whether it was safe to do word <B>OUT</B>s (<B>OUT DX,AX</B>) to set Index/Data register pairs in a single instruction. Long ago, there were a few computers with buses that weren&#146;t quite PC-compatatible, in that the two bytes in each word <B>OUT</B> went to the VGA in the wrong order: Data register first, then Index register, with predictably disastrous results. Consequently, I generally wrote my code in those days to use two 8-bit <B>OUT</B>s to set indexed registers. Later on, I made it a habit to use macros that could do either one 16-bit <B>OUT</B> or two 8-bit <B>OUT</B>s, depending on how I chose to assemble the code, and in fact you&#146;ll find both ways of dealing with <B>OUT</B>s sprinkled through the code in this part of the book. Using macros for word OUTs is still not a bad idea in that it does no harm, but in my opinion it&#146;s no longer necessary. Word <B>OUT</B>s are standard now, and it&#146;s been a long time since I&#146;ve heard of them causing any problems.</P><P><BR></P>

View file

@ -38,7 +38,7 @@
<P><BR></P>
<P>The key to understanding Listing 26.1 is understanding the effect of ANDing the rotated CPU data with the contents of the Bit Mask register. The CPU data is the pattern for the character to be drawn, with bits equal to 1 indicating where character pixels are to appear. The Data Rotate register is set to rotate the CPU data to pixel-align it, since without rotation characters could only be drawn on byte boundaries.
</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/26-01i.jpg"><TD WIDTH="95%"><SMALL><I>As I pointed out in Chapter 25, the CPU is perfectly capable of rotating the data itself, and it&#146;s often the case that that&#146;s more efficient. The problem with using the Data Rotate register is that the <B>OUT</B> that sets that register is time-consuming, especially for proportional text, which requires a different rotation for each character. Also, if the code performs full-byte accesses to display memory&#151;that is, if it combines pieces of two adjacent characters into one byte&#151;whenever possible for efficiency, the CPU generally has to do extra work to prepare the data so the VGA&#146;s rotator can handle it.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>As I pointed out in Chapter 25, the CPU is perfectly capable of rotating the data itself, and it&#146;s often the case that that&#146;s more efficient. The problem with using the Data Rotate register is that the <B>OUT</B> that sets that register is time-consuming, especially for proportional text, which requires a different rotation for each character. Also, if the code performs full-byte accesses to display memory&#151;that is, if it combines pieces of two adjacent characters into one byte&#151;whenever possible for efficiency, the CPU generally has to do extra work to prepare the data so the VGA&#146;s rotator can handle it.</I></SMALL>
</TABLE>
<P>At the same time that the Data Rotate register is set, the Bit Mask register is set to allow the CPU to modify only that portion of the display memory byte accessed that the pixel-aligned character falls in, so that other characters and/or graphics data won&#146;t be wiped out. The result of ANDing the rotated CPU data byte with the contents of the Bit Mask register is a bit mask that allows only the bits equal to 1 in the original character pattern (rotated and masked to provide pixel alignment) to be modified by the CPU; all other bits come straight from the latches. The latches should have previously been loaded from the target address, so the effect of the ultimate synthesized bit mask value is to allow the CPU to modify only those pixels in display memory that correspond to the 1 bits in that part of the pixel-aligned character that falls in the currently addressed byte. The color of the pixels set by the CPU is determined by the contents of the Set/Reset register.
</P>

View file

@ -55,7 +55,7 @@
<BR><A HREF="javascript:displayWindow('images/27-01.jpg',409,401)"> --><FONT COLOR="#000077"><B>Figure 27.1</B></FONT></A>&nbsp;&nbsp;<I>VGA data flow in write mode 2.</I>
</P>
<TABLE WIDTH="100%">
<TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/27-01i.jpg"><TD WIDTH="95%"><SMALL><I>It&#146;s worth noting two differences between write mode 2 and write mode 0, the standard write mode of the VGA. First, rotation of the CPU data byte does not take place in write mode 2. Second, the Set/Reset and Enable Set/Reset registers have no effect in write mode 2.</I></SMALL>
<TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>It&#146;s worth noting two differences between write mode 2 and write mode 0, the standard write mode of the VGA. First, rotation of the CPU data byte does not take place in write mode 2. Second, the Set/Reset and Enable Set/Reset registers have no effect in write mode 2.</I></SMALL>
</TABLE>
<P>Now that we understand the mechanics of write mode 2, we can step back and get a feel for what it might be useful for. View bits 3-0 of the CPU byte as a single pixel in one of 16 colors. Next imagine that nibble turned sideways and written across the four planes, one bit to a plane. Finally, expand each of the bits to a byte, as shown in Figure 27.2, so that 8 pixels are drawn in the color selected by bits 3-0 of the CPU byte. Within the constraints of the VGA&#146;s data paths, that&#146;s exactly what write mode 2 does.
</P>

View file

@ -38,7 +38,7 @@
<P><BR></P>
<P>&#147;That&#146;s an interesting application of write mode 2,&#148; you may well say, &#147;but is it really useful?&#148; While the ability to convert chunky bitmaps into VGA bitmaps does have its uses, Listing 27.1 is primarily intended to illustrate the mechanics of write mode 2.
</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/27-02i.jpg"><TD WIDTH="95%"><SMALL><I>For performance, it&#146;s best to store 16-color bitmaps in pre-separated four-plane format in system memory, and copy one plane at a time to the screen. Ideally, such bitmaps should be copied one scan line at a time, with all four planes completed for one scan line before moving on to the next. I say this because when entire images are copied one plane at a time, nasty transient color effects can occur as one plane becomes visibly changed before other planes have been modified.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>For performance, it&#146;s best to store 16-color bitmaps in pre-separated four-plane format in system memory, and copy one plane at a time to the screen. Ideally, such bitmaps should be copied one scan line at a time, with all four planes completed for one scan line before moving on to the next. I say this because when entire images are copied one plane at a time, nasty transient color effects can occur as one plane becomes visibly changed before other planes have been modified.</I></SMALL>
</TABLE>
<H4 ALIGN="LEFT"><A NAME="Heading6"></A><FONT COLOR="#000077">Drawing Color-Patterned Lines Using Write Mode 2</FONT></H4>
<P>A more serviceable use of write mode 2 is shown in the program presented in Listing 27.2. The program draws multicolored horizontal, vertical, and diagonal lines, basing the color patterns on passed color tables. Write mode 2 is ideal because in this application color can vary from one pixel to the next, and in write mode 2 all that&#146;s required to set pixel color is a change of the lower nibble of the byte written by the CPU. Set/reset could be used to achieve the same result, but an index/data pair of <B>OUT</B>s would be required to set the Set/Reset register to each new color. Similarly, the Map Mask register could be used in write mode 0 to set pixel color, but in this case not only would an index/data pair of <B>OUT</B>s be required but there would also be no guarantee that data already in display memory wouldn&#146;t interfere with the color of the pixel being drawn, since the Map Mask register allows only selected planes to be drawn to.</P>

View file

@ -39,7 +39,7 @@
<P>By the way, the code in Listing 28.1 is intended only to illustrate read mode 0, and is, in general, a poor way to perform animation, since it&#146;s slow and tends to flicker. Later in this book, we&#146;ll take a look at some far better VGA animation techniques.
</P>
<P>As you&#146;d expect, neither the read mode nor the setting of the Read Map register affects CPU <I>writes</I> to VGA memory in any way.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/28-01i.jpg"><TD WIDTH="95%"><SMALL><I>An important point regarding reading VGA memory involves the VGA&#146;s latches. (Remember that each of the four latches stores a byte for one plane; on CPU writes, the latches can provide some or all of the data written to display memory, allowing fast copying and efficient pixel masking.) Whenever the CPU reads a given address in VGA memory, each of the four latches is loaded with the contents of the byte at that address in its respective plane. Even though the CPU only receives data from one plane in read mode 0, all four planes are always read, and the values read are stored in the latches. This is true in read mode 1 as well. In short, whenever the CPU reads VGA memory in any read mode, all four planes are read and all four latches are always loaded.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>An important point regarding reading VGA memory involves the VGA&#146;s latches. (Remember that each of the four latches stores a byte for one plane; on CPU writes, the latches can provide some or all of the data written to display memory, allowing fast copying and efficient pixel masking.) Whenever the CPU reads a given address in VGA memory, each of the four latches is loaded with the contents of the byte at that address in its respective plane. Even though the CPU only receives data from one plane in read mode 0, all four planes are always read, and the values read are stored in the latches. This is true in read mode 1 as well. In short, whenever the CPU reads VGA memory in any read mode, all four planes are read and all four latches are always loaded.</I></SMALL>
</TABLE>
<H3><A NAME="Heading4"></A><FONT COLOR="#000077">Read Mode 1</FONT></H3>
<P>Read mode 0 is the workhorse read mode, but it&#146;s got an annoying limitation: Whenever you want to determine the color of a given pixel in read mode 0, you have to perform four VGA memory reads, one for each plane, and then interpret the four bytes you&#146;ve read as eight 16-color pixels. That&#146;s a lot of programming. The code is also likely to run slowly, all the more so because a standard IBM VGA takes an average of 1.1 microseconds to complete each memory read, and read mode 0 requires four reads in order to read the four planes, not to mention the even greater amount of time taken by the <B>OUT</B>s required to switch between the planes. (1.1 microseconds may not sound like much, but on a 66-MHz 486, it&#146;s 73 clock cycles! Local-bus VGAs can be a good deal faster, but a read from the fastest local-bus adapter I&#146;ve yet seen would still cost in the neighborhood of 10 486/66 cycles.)</P>

View file

@ -37,7 +37,7 @@
</CENTER>
<P><BR></P>
<TABLE WIDTH="100%">
<TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/29-01i.jpg"><TD WIDTH="95%"><SMALL><I>While these requirements are no problem if you&#146;re simply calling a subroutine in order to save an image from your program, they pose a considerable problem if you&#146;re designing a hot-key operated TSR that can capture a screen image at any time. With the EGA specifically, there&#146;s never any way to tell what state the registers are currently in, since the registers aren&#146;t readable. (More on this issue later in this chapter.) As a result, any TSR that sets the Bit Mask to 0FFH, the Data Rotate register to 0, and so on runs the risk of interfering with the drawing code of the program that&#146;s already running.</I></SMALL>
<TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>While these requirements are no problem if you&#146;re simply calling a subroutine in order to save an image from your program, they pose a considerable problem if you&#146;re designing a hot-key operated TSR that can capture a screen image at any time. With the EGA specifically, there&#146;s never any way to tell what state the registers are currently in, since the registers aren&#146;t readable. (More on this issue later in this chapter.) As a result, any TSR that sets the Bit Mask to 0FFH, the Data Rotate register to 0, and so on runs the risk of interfering with the drawing code of the program that&#146;s already running.</I></SMALL>
</TABLE>
<P>What&#146;s the solution? Frankly, the solution is to get VGA-specific. A TSR designed for the VGA can simply read out and save the state of the registers of interest, program those registers as needed, save the screen image, and restore the original settings. From a programmer&#146;s perspective, readable registers are certainly near the top of the list of things to like about the VGA! The remaining installed base of EGAs is steadily dwindling, and you may be able to ignore it as a market today, as you couldn&#146;t even a year or two ago.
</P>
@ -51,7 +51,7 @@ DISPLAYED_SCREEN_SIZEequ(640/8)*480
</P>
<P>Finally, be aware that the screen capture and restore programs in Listings 29.1 and 29.2 are only appropriate for EGA/VGA modes 0DH, 0EH, 0FH, 010H, and 012H, since they assume a fourconfiguration of EGA/VGA memory. In all text modes and in CGA graphics modes, and in VGA modes 11H and 13H as well, display memory can simply be written to disk and read back as a linear block of memory, just like a normal array.</P>
<P>While Listings 29.1 and 29.2 are written in assembly, the principles they illustrate apply equally well to high-level languages. In fact, there&#146;s no need for any assembly at all when saving an EGA/VGA screen, as long as the high-level language you&#146;re using can perform direct port I/O to set up the adapter and can read and write display memory directly.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/29-02i.jpg"><TD WIDTH="95%"><SMALL><I>One tip if you&#146;re saving and restoring the screen from a high-level language on an EGA, though: After you&#146;ve completed the save or restore operation, be sure to put any registers that you&#146;ve changed back to their default settings. Some high-level languages (and the BIOS as well) assume that various registers are left in a certain state, so on the EGA it&#146;s safest to leave the registers in their most likely state. On the VGA, of course, you can just read the registers out before you change them, then put them back the way you found them when you&#146;re done.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>One tip if you&#146;re saving and restoring the screen from a high-level language on an EGA, though: After you&#146;ve completed the save or restore operation, be sure to put any registers that you&#146;ve changed back to their default settings. Some high-level languages (and the BIOS as well) assume that various registers are left in a certain state, so on the EGA it&#146;s safest to leave the registers in their most likely state. On the VGA, of course, you can just read the registers out before you change them, then put them back the way you found them when you&#146;re done.</I></SMALL>
</TABLE>
<H3><A NAME="Heading4"></A><FONT COLOR="#000077">16 Colors out of 64</FONT></H3>
<P>How does one produce the 64 colors from which the 16 colors displayed by the EGA can be chosen? The answer is simple enough: There&#146;s a BIOS function that lets you select the mapping of the 16 possible pixel values to the 64 possible colors. Let&#146;s lay out a bit of background before proceeding, however.

View file

@ -39,7 +39,7 @@
<H3><A NAME="Heading5"></A><FONT COLOR="#000077">Overscan</FONT></H3>
<P>While we&#146;re at it, I&#146;m going to touch on overscan. Overscan is the color of the border of the display, the rectangular area around the edge of the monitor that&#146;s outside the region displaying active video data but inside the blanking area. The overscan (or border) color can be programmed to any of the 64 possible colors by either setting Attribute Controller register 11H directly or calling video function 10H, subfunction 1.
</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/29-03i.jpg"><TD WIDTH="95%"><SMALL><I>On ECD-compatible monitors, however, there&#146;s too little scan time to display a proper border when the EGA is in 350-scan-line mode, so overscan should always be 0 (black) unless you&#146;re in 200-scanmode. Note, though, that a VGA can easily display a border on a VGA-compatible monitor, and VGAs are in fact programmed at mode set for an 8-pixel-wide border in all modes; all you need do is set the overscan color on any VGA to see the border.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>On ECD-compatible monitors, however, there&#146;s too little scan time to display a proper border when the EGA is in 350-scan-line mode, so overscan should always be 0 (black) unless you&#146;re in 200-scanmode. Note, though, that a VGA can easily display a border on a VGA-compatible monitor, and VGAs are in fact programmed at mode set for an 8-pixel-wide border in all modes; all you need do is set the overscan color on any VGA to see the border.</I></SMALL>
</TABLE>
<H3><A NAME="Heading6"></A><FONT COLOR="#000077">A Bonus Blanker</FONT></H3>
<P>An interesting bonus: The Attribute Controller provides a very convenient way to blank the screen, in the form of the aforementioned bit 5 of the Attribute Controller Index register (at address 3C0H after the Input Status 1 register&#151;3DAH in color, 3BAH in monochrome&#151;has been read and on every other write to 3C0H thereafter). Whenever bit 5 of the AC Index register is 0, video data is cut off, effectively blanking the screen. Setting bit 5 of the AC Index back to 1 restores video data immediately. Listing 29.4 illustrates this simple but effective form of screen blanking.

View file

@ -44,7 +44,7 @@
<P>Setting the split-screen-related registers is not as simple a matter as merely outputting the right values to the right registers; timing is also important. The split screen start scan line value is checked against the number of each scan line as that scan line is displayed, which means that the split screen start scan line potentially takes effect the moment it is set. In other words, if the screen is displaying scan line 15 and you set the split screen start to 16, that change will be picked up immediately and the split screen will start after the next scan line. This is markedly different from changes to the start address, which take effect only at the start of the next frame.
</P>
<P>The instantly-effective nature of the split screen is a bit of a problem, not because the changed screen appears as soon as the new split screen start scan line is set&#151;that seems to me to be an advantage&#151;but because the changed screen can appear <I>before</I> the new split screen start scan line is set.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/30-01i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Remember, the split screen start scan line is spread out over two or three registers. What if the incompletely-changed value matches the current scan line after you&#146;ve set one register but before you&#146;ve set the rest? For one frame, you&#146;ll see the split screen in a wrong place&#151;possibly a very wrong place&#151;resulting in jumping and flicker.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Remember, the split screen start scan line is spread out over two or three registers. What if the incompletely-changed value matches the current scan line after you&#146;ve set one register but before you&#146;ve set the rest? For one frame, you&#146;ll see the split screen in a wrong place&#151;possibly a very wrong place&#151;resulting in jumping and flicker.</I></SMALL>
</TABLE>
<P>The solution is simple: Set the split screen start scan line at a time when it can&#146;t possibly match the currently displayed scan line. The easy way to do that is to set it when there isn&#146;t any currently displayed scan line&#151;during vertical non-display time. One safe time that&#146;s easy to find is the start of the vertical sync pulse, which is typically pretty near the middle of vertical non-display time, and that&#146;s the approach I&#146;ve followed in Listing 30.1. I&#146;ve also disabled interrupts during the period when the split screen registers are being set. This isn&#146;t absolutely necessary, but if it&#146;s not done, there&#146;s the possibility that an interrupt will occur between register sets and delay the later register sets until display time, again causing flicker.
</P>
@ -54,7 +54,7 @@
</P>
<P>The bug is this: The first scan line of the EGA split screen&#151;the scan line starting at offset zero in display memory&#151;is displayed not once but twice. In other words, the first line of split screen display memory, and only the first line, is replicated one unnecessary time, pushing all the other lines down by one.</P>
<P>That&#146;s not a fatal bug, of course. In fact, if the first few scan lines are identical, it&#146;s not even noticeable. The EGA&#146;s split-screen bug can produce visible distortion given certain patterns, however, so you should try to make the top few lines identical (if possible) when designing split-screen images that might be displayed on EGAs, and you should in any case check how your split-screens look on both VGAs and EGAs.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/30-02i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>I have an important caution here: Don&#146;t count on the EGA&#146;s split-screen bug; that is, don&#146;t rely on the first scan line being doubled when you design your split screens. IBM designed and made the original EGA, but a lot of companies cloned it, and there&#146;s no guarantee that all EGA clones copy the bug. It is a certainty, at least, that the VGA didn&#146;t copy it.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>I have an important caution here: Don&#146;t count on the EGA&#146;s split-screen bug; that is, don&#146;t rely on the first scan line being doubled when you design your split screens. IBM designed and made the original EGA, but a lot of companies cloned it, and there&#146;s no guarantee that all EGA clones copy the bug. It is a certainty, at least, that the VGA didn&#146;t copy it.</I></SMALL>
</TABLE>
<P>There&#146;s another respect in which the EGA is inferior to the VGA when it comes to the split screen, and that&#146;s in the area of panning when the split screen is on. This isn&#146;t a bug&#151;it&#146;s just one of the many areas in which the VGA&#146;s designers learned from the shortcomings of the EGA and went the EGA one better.
</P><P><BR></P>

View file

@ -41,7 +41,7 @@
</P>
<P>Horizontal smooth panning works just fine, although I&#146;ve always harbored some doubts that any one horizontal-smooth-panning approach works properly on all display board clones. (More on this later.) There&#146;s a catch when using horizontal smooth panning with the split screen up, though, and it&#146;s a serious catch: You can&#146;t byte-pan the split screen (which always starts at offset zero, no matter what the setting of the start address registers)&#151;but you <I>can</I> pel-pan the split screen.</P>
<P>Put another way, when the normal portion of the screen is horizontally smooth-panned, the split screen portion moves a pixel at a time until it&#146;s time to move to the next byte, then jumps back to the start of the current byte. As the top part of the screen moves smoothly about, the split screen will move and jump, move and jump, over and over. Believe me, it&#146;s not a pretty sight.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/30-03i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>What&#146;s to be done? On the EGA, nothing. Unless you&#146;re willing to have your users&#146; eyes doing the jitterbug, don&#146;t use horizontal smooth scrolling while the split screen is up. Byte panning is fine&#151;just don&#146;t change the Pel Panning register from its default setting.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>What&#146;s to be done? On the EGA, nothing. Unless you&#146;re willing to have your users&#146; eyes doing the jitterbug, don&#146;t use horizontal smooth scrolling while the split screen is up. Byte panning is fine&#151;just don&#146;t change the Pel Panning register from its default setting.</I></SMALL>
</TABLE>
<P>On the VGA, there is recourse. A VGA-only bit, bit 5 of the AC Mode Control register (AC register 10H), turns off pel panning in the split screen. In other words, when this bit is set to 1, pel panning is reset to zero before the first line of the split screen, and remains zero until the end of the frame. This doesn&#146;t allow you to pan the split screen horizontally, mind you&#151;there&#146;s no way to do that&#151;but it does let you pan the normal screen while the split screen stays rock-solid. This can be used to produce an attractive &#147;streaming tape&#148; effect in the normal screen while the split screen is used to display non-moving information.
</P>

View file

@ -44,7 +44,7 @@
<P>You see, the start address is loaded into the EGA&#146;s or VGA&#146;s internal display memory pointer once per frame. The internal pointer is then advanced, byte-by-byte and line-by-line, until the end of the frame (with a possible resetting to zero if the split screen line is reached), and is then reloaded for the next frame. That&#146;s straightforward enough; the real question is, <I>Exactly when is the start address loaded?</I></P>
<P>In his excellent book <I>Programmer&#146;s Guide to PC Video Systems</I> (Microsoft Press) Richard Wilton says that the start address is loaded at the start of the vertical sync pulse. (Wilton calls it vertical retrace, which can also be taken to mean vertical non-display time, but given that he&#146;s testing the vertical sync status bit in the Input Status 0 register, I assume he means that the start address is loaded at the start of vertical sync.) Consequently, he waits until the <I>end</I> of the vertical sync pulse to set the start address registers, confident that the start address won&#146;t take effect until the next frame.</P>
<P>I&#146;m sure Richard is right when it comes to the real McCoy IBM VGA and EGA, but I&#146;m less confident that every clone out there loads the start address at the start of vertical sync.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/30-04i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>For that very reason, I generally advise people not to use horizontal smooth panning unless they can test their software on all the makes of display adapter it might run on. I&#146;ve used Richard&#146;s approach in Listings 30.1 and 30.2, and so far as I&#146;ve seen it works fine, but be aware that there are potential, albeit unproven, hazards to relying on the setting of the start address registers to occur at a specific time in the frame.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>For that very reason, I generally advise people not to use horizontal smooth panning unless they can test their software on all the makes of display adapter it might run on. I&#146;ve used Richard&#146;s approach in Listings 30.1 and 30.2, and so far as I&#146;ve seen it works fine, but be aware that there are potential, albeit unproven, hazards to relying on the setting of the start address registers to occur at a specific time in the frame.</I></SMALL>
</TABLE>
<P>The interaction of the start address registers and the Pel Panning register is worthy of note. After waiting for the end of vertical sync to set the start address in Listing 30.2, I wait for the start of the <I>next</I> vertical sync to set the Pel Panning register. That&#146;s because the start address doesn&#146;t take effect until the start of the next frame, but the pel panning setting takes effect at the start of the next line; if we set the pel panning at the same time we set the start address, we&#146;d get a whole frame with the old start address and the new pel panning settings mixed together, causing the screen to jump. As with the split screen registers, it&#146;s safest to set the Pel Panning register during non-display time. For maximum reliability, we&#146;d have interrupts off from the time we set the start address registers to the time we change the pel planning setting, to make sure an interrupt doesn&#146;t come in and cause us to miss the start of a vertical sync and thus get a mismatched pel panning/start address pair for a frame, although for modularity I haven&#146;t done this in Listing 30.2. (Also, doing so would require disabling interrupts for much too long a time.)</P>
<P>What if you wanted to pan faster? Well, you could of course just move two pixels at a time rather than one; I assure you no one will ever notice when you&#146;re panning at a rate of 10 or more times per second.</P><P><BR></P>

View file

@ -243,7 +243,7 @@ void main()
<!-- END CODE //-->
<P>The first thing you&#146;ll notice when you run this code is that the speed of 360&#215;480 256-color mode is pretty good, especially considering that most of the program is im-plemented in C.
</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/32-01i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>Drawing in 360&#215;480 256-color mode can sometimes actually be faster than in the 16-color modes, because the byte-per-pixel display memory organization of 256-color mode eliminates the need to read display memory before writing to it in order to isolate individual pixels coexisting within a single byte. In addition, 360&#215;480 256-color mode is a variant of Mode X, which we&#146;ll encounter in detail in Chapter 47, and supports all the high-performance features of Mode X.</I></SMALL>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>Drawing in 360&#215;480 256-color mode can sometimes actually be faster than in the 16-color modes, because the byte-per-pixel display memory organization of 256-color mode eliminates the need to read display memory before writing to it in order to isolate individual pixels coexisting within a single byte. In addition, 360&#215;480 256-color mode is a variant of Mode X, which we&#146;ll encounter in detail in Chapter 47, and supports all the high-performance features of Mode X.</I></SMALL>
</TABLE>
<P><BR></P>
<CENTER>

View file

@ -40,7 +40,7 @@
<P>The DAC can be loaded either directly or through subfunctions 10H (for a single DAC register) or 12H (for a block of DAC registers) of the BIOS video service interrupt 10H, function 10H, described in Chapter 33. For cycling the contents of the entire DAC, the block-load function (invoked by executing <B>INT</B> 10H with AH = 10H and AL = 12H to load a block of CX DAC locations, starting at location BX, from the block of RGB triplets&#151;3 bytes per triplet&#151;starting at ES:DX into the DAC) would be the better of the two, due to the considerably greater efficiency of calling the BIOS once rather than 256 times. At any rate, we&#146;d like to use one or the other of the BIOS functions for color cycling, because we know that whenever possible, one should use a BIOS function in preference to accessing hardware directly, in the interests of avoiding compatibility problems. In the case of color cycling, however, it is emphatically <I>not</I> possible to use either of the BIOS functions, for they have problems. Serious problems.</P>
<P>The difficulty is this: IBM&#146;s BIOS specification describes exactly how the parameters passed to the BIOS control the loading of DAC locations, and all clone BIOSes meet that specification scrupulously, which is to say that if you invoke <B>INT</B> 10H, function 10H, subfunction 12H with a given set of parameters, you can be sure that you will end up with the same values loaded into the same DAC locations on all VGAs from all vendors. IBM&#146;s spec does <I>not</I>, however, describe whether vertical retrace should be waited for before loading the DAC, nor does it mention whether video should be left enabled while loading the DAC, leaving cloners to choose whatever approach they desire&#151;and, alas, every VGA cloner seems to have selected a different approach.</P>
<P>I tested four clone VGAs from different manufacturers, some in a 20 MHz 386 machine and some in a 10 MHz 286 machine. Two of the four waited for vertical retrace before loading the DAC; two didn&#146;t. Two of the four blanked the display while loading the DAC, resulting in flickering bars across the screen. One showed speckled pixels spattered across the top of the screen while the DAC was being loaded. Also, not one was able to load all 256 DAC locations without showing <I>some</I> sort of garbage on the screen for at least one frame, but that&#146;s not the BIOS&#146;s fault; it&#146;s a problem endemic to the VGA.</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/34-01i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>These findings lead me inexorably to the conclusion that the BIOS should not be used to load the DAC dynamically. That is, if you&#146;re loading the DAC just once in preparation for a graphics session&#151;sort of a DAC mode set&#151;by all means load by way of the BIOS. No one will care that some garbage is displayed for a single frame; heck, I have boards that bounce and flicker and show garbage every time I do a mode set, and the amount of garbage produced by loading the DAC once is far less noticeable. If, however, you intend to load the DAC repeatedly for color cycling, avoid the BIOS DAC load functions like the plague. They will bring you only heartache.</I></SMALL>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>These findings lead me inexorably to the conclusion that the BIOS should not be used to load the DAC dynamically. That is, if you&#146;re loading the DAC just once in preparation for a graphics session&#151;sort of a DAC mode set&#151;by all means load by way of the BIOS. No one will care that some garbage is displayed for a single frame; heck, I have boards that bounce and flicker and show garbage every time I do a mode set, and the amount of garbage produced by loading the DAC once is far less noticeable. If, however, you intend to load the DAC repeatedly for color cycling, avoid the BIOS DAC load functions like the plague. They will bring you only heartache.</I></SMALL>
</TABLE>
<P>As but one example of the unsuitability of the BIOS DAC-loading functions for color cycling, imagine that you want to cycle all 256 colors 70 times a second, which is once per frame. In order to accomplish that, you would normally wait for the start of the vertical sync signal (marking the end of the frame), then call the BIOS to load the DAC. On some boards&#151;boards with BIOSes that don&#146;t wait for vertical sync before loading the DAC&#151;that will work pretty well; you will, in fact, load the DAC once a frame. On other boards, however, it will work very poorly indeed; your program will wait for the start of vertical sync, and then the BIOS will wait for the start of the next vertical sync, with the result being that the DAC gets loaded only once every <I>two</I> frames. Sadly, there&#146;s no way, short of actually profiling the performance of BIOS DAC loads, for you to know which sort of BIOS is installed in a particular computer, so unless you can always control the brand of VGA your software will run on, you really can&#146;t afford to color cycle by calling the BIOS.</P>
<P>Which is not to say that loading the DAC directly is a picnic either, as we&#146;ll see next.</P>

View file

@ -51,7 +51,7 @@
<P>First of all, I&#146;d like to point out that when color cycling does work, it&#146;s a thing of beauty. Assemble Listing 34.1 so that it doesn&#146;t use the BIOS to load the DAC, doesn&#146;t guard against interrupts, and uses 286-specific instructions if your computer supports them. Then tinker with <B>CYCLE_SIZE</B> until the color cycling is perfectly clean on your computer. Color cycling looks stunningly smooth, doesn&#146;t it? And this is crude color cycling, working with the default color set; switch over to a color set that gradually works its way through various hues and saturations, and you could get something that looks for all the world like true-color animation (albeit working with a small subset of the full spectrum at any one time).</P>
<P>Given that, how can we take advantage of color cycling within the limitations of loading the DAC? The simplest approach, and my personal favorite, is that of cycling a portion of the DAC while using the rest of the DAC locations for other, non-cycling purposes. For example, you might allocate 32 DAC locations to the aforementioned sunset, reserve 160 additional locations for use in drawing a static mountain scene, and employ the remaining 64 locations to draw images of planes, cars, and the like in the foreground. The 32 sunset colors could be cycled cleanly, and the other 224 colors would remain the same throughout the program, or would change only occasionally.</P>
<P>That suggests a second possibility: If you have several different color sets to be cycled, interleave the loading so that only one color set is cycled per frame. Suppose you are animating a night scene, with stars twinkling in the background, meteors streaking across the sky, and a spaceship moving across the screen with its jets flaring. One way to produce most of the necessary effects with little effort would be to draw the stars in several attributes and then cycle the colors for <I>those</I> attributes, draw the meteor paths in successive attributes, one for each pixel, and then cycle the colors for those attributes, and do much the same for the jets. The only remaining task would be to animate the spaceship across the screen, which is not a particularly difficult task.</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/34-02i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>The key to getting all the color cycling to work in the above example, however, would be to assign each color cycling task a different part of the DAC, with each part cycled independently as needed. If, as is likely, the total number of DAC locations cycled proved to be too great to manage in one frame, you could simply cycle the colors of the stars after one frame, the colors of the meteors after the next, and the colors of the jets after yet another frame, then back around to cycling the colors of the stars. By splitting up the DAC in this manner and interleaving the cycling tasks, you can perform a great deal of seemingly complex color animation without loading very much of the DAC during any one frame.</I></SMALL>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>The key to getting all the color cycling to work in the above example, however, would be to assign each color cycling task a different part of the DAC, with each part cycled independently as needed. If, as is likely, the total number of DAC locations cycled proved to be too great to manage in one frame, you could simply cycle the colors of the stars after one frame, the colors of the meteors after the next, and the colors of the jets after yet another frame, then back around to cycling the colors of the stars. By splitting up the DAC in this manner and interleaving the cycling tasks, you can perform a great deal of seemingly complex color animation without loading very much of the DAC during any one frame.</I></SMALL>
</TABLE>
<P><BR></P>
<CENTER>

View file

@ -127,7 +127,7 @@ void main()
<H4 ALIGN="LEFT"><A NAME="Heading7"></A><FONT COLOR="#000077">Looking at EVGALine</FONT></H4>
<P>The <B>EVGALine</B> function itself performs four operations. <B>EVGALine</B> first sets up the VGA&#146;s hardware so that all pixels drawn will be in the desired color. This is accomplished by setting two of the VGA&#146;s registers, the Enable Set/Reset register and the Set/Reset register. Setting the Enable Set/Reset to the value 0FH, as is done in <B>EVGALine</B>, causes all drawing to produce pixels in the color contained in the Set/Reset register. Setting the Set/Reset register to the passed color, in conjunction with the Enable Set/Reset setting of 0FH, causes all drawing done by <B>EVGALine</B> and the functions it calls to generate the passed color. In summary, setting up the Enable Set/Reset and Set/Reset registers in this way causes the remainder of <B>EVGALine</B> to draw a line in the specified color.</P>
<P><B>EVGALine</B> next performs a simple check to cut in half the number of line orientations that must be handled separately. Figure 35.4 shows the eight possible line orientations among which a Bresenham&#146;s algorithm implementation must distinguish. (In interpreting Figure 35.4, assume that lines radiate outward from the center of the figure, falling into one of eight octants delineated by the horizontal and vertical axes and the two diagonals.) The need to categorize lines into these octants falls out of the major/minor axis nature of the algorithm; the orientations are distinguished by which coordinate forms the major axis and by whether each of X and Y increases or decreases from the line start to the line end.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/35-01i.jpg"><TD WIDTH="95%"><SMALL><I>A moment of thought will show, however, that four of the line orientations are redundant. Each of the four orientations for which <B>DeltaY</B>, the Y component of the line, is less than 0 (that is, for which the line start Y coordinate is greater than the line end Y coordinate) can be transformed into one of the four orientations for which the line start Y coordinate is less than the line end Y coordinate simply by reversing the line start and end coordinates, so that the line is drawn in the other direction. <B>EVGALine</B> does this by swapping (X0,Y0) (the line start coordinates) with (X1,Y1) (the line end coordinates) whenever Y0 is greater than Y1.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>A moment of thought will show, however, that four of the line orientations are redundant. Each of the four orientations for which <B>DeltaY</B>, the Y component of the line, is less than 0 (that is, for which the line start Y coordinate is greater than the line end Y coordinate) can be transformed into one of the four orientations for which the line start Y coordinate is less than the line end Y coordinate simply by reversing the line start and end coordinates, so that the line is drawn in the other direction. <B>EVGALine</B> does this by swapping (X0,Y0) (the line start coordinates) with (X1,Y1) (the line end coordinates) whenever Y0 is greater than Y1.</I></SMALL>
</TABLE>
<P>This accomplished, <B>EVGALine</B> must still distinguish among the four remaining line orientations. Those four orientations form two major categories, orientations for which the X dimension is the major axis of the line and orientations for which the Y dimension is the major axis. As shown in Figure 35.4, octants 1 (where X increases from start to finish) and 2 (where X decreases from start to finish) fall into the latter category, and differ in only one respect, the direction in which the X coordinate moves when it changes. Handling of the running error of the line is exactly the same for both cases, as one would expect given the symmetry of lines differing only in the sign of <B>DeltaX</B>, the X coordinate of the line. Consequently, for those cases where <B>DeltaX</B> is less than zero, the direction of X movement is made negative, and the absolute value of <B>DeltaX</B> is used for error term calculations.</P>
<P>Similarly, octants 0 (where X increases from start to finish) and 3 (where X decreases from start to finish) differ only in the direction in which the X coordinate moves when it changes. The difference between line drawing in octants 0 and 3 and line drawing in octants 1 and 2 is that in octants 0 and 3, since X is the major axis, the X coordinate changes on every pixel of the line and the Y coordinate changes only when the running error of the line dictates. In octants 1 and 2, the Y coordinate changes on every pixel and the X coordinate changes only when the running error dictates, since Y is the major axis.</P>

View file

@ -41,7 +41,7 @@
<P>When I went to speed up run-length slice lines, I initially manually converted the last chapter&#146;s C code into assembly. Then I streamlined the register usage and used <B>REP STOS</B> wherever possible. Listing 37.1 is that code. At that point, line drawing was surely faster, although I didn&#146;t know exactly how much faster. Equally surely, there were significant optimizations yet to be made, and I was itching to get on to them, for they were bound to be a lot more interesting than a basic C-to-assembly port.</P>
<P>Ego intervened at this point, however. I wanted to know how much of a speed-up I had already gotten, so I timed the performance of the C code and compared it to the assembly code. To my horror, I found that I had not gotten even a two-times improvement! I couldn&#146;t understand how that could be&#151;the C code was decidedly unoptimized&#151;until I hit on the idea of measuring the maximum memory speed of the VGA to which I was drawing.</P>
<P>Bingo. The Paradise VGA in my 486/33 is fast for a single display-memory write, because it buffers the data, lets the CPU go on its merry way, and finishes the write when display memory is ready. However, the maximum rate at which data can be written to the adapter turns out to be no more than one byte every microsecond. Put another way, you can only write one byte to this adapter every 33 clock cycles on a 486/33. Therefore, no matter how fast I made the line-drawing code, it could never draw more than 1,000,000 pixels per second in 256-color mode in my system. The C code was already drawing at about half that rate, so the potential speed-up for the assembly code was limited to a maximum of two times, which is pretty close to what Listing 37.1 did, in fact, achieve. When I compared the C and assembly implementations drawing to normal system (nondisplay) memory, I found that the assembly code was actually four times as fast as the C code.</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/37-01i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>In fact, Listing 37.1 draws VGA lines at about 92 percent of the maximum possible rate in my system&#151;that is, it draws very nearly as fast as the VGA hardware will allow. All the optimization in the world would get me less than 10 percent faster line drawing&#151;and only if I eliminated all overhead, an unlikely proposition at best. The code isn&#146;t fully optimized, but so what?</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>In fact, Listing 37.1 draws VGA lines at about 92 percent of the maximum possible rate in my system&#151;that is, it draws very nearly as fast as the VGA hardware will allow. All the optimization in the world would get me less than 10 percent faster line drawing&#151;and only if I eliminated all overhead, an unlikely proposition at best. The code isn&#146;t fully optimized, but so what?</I></SMALL>
</TABLE>
<P>Now it&#146;s true that faster line-drawing code would likely be more beneficial on faster VGAs, especially local-bus VGAs, and in slower systems. For that reason, I&#146;ll list a variety of potential optimizations to Listing 37.1. On the other hand, it&#146;s also true that Listing 37.1 is capable of drawing lines at a rate of 2.2 million pixels per second on a 486/ 33, given fast enough VGA memory, so it should be able to drive almost any non-local-bus VGA at nearly full speed. In short, Listing 37.1 is very fast, and, in many systems, further optimization is basically a waste of time.
</P>

View file

@ -49,7 +49,7 @@
<P>All edges of a polygon except those that are flat tops or flat bottoms will be considered either right edges or left edges, regardless of slope. The left edge is the one that starts with the leftmost line down from the top of the polygon.
</P>
<P>These rules ensure that no pixel is drawn more than once when adjacent polygons are filled, and that if polygons cover the full 360-degree range around a pixel, then that pixel will be drawn once and only once&#151;just what we need in order to be able to fit filled polygons together seamlessly.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/38-01i.jpg"><TD WIDTH="95%" VALIGN="TOP" ALIGN="LEFT"><SMALL><I>This sort of non-overlapping polygon filling isn&#146;t ideal for all purposes. Polygons are skewed toward the top and left edges, which not only introduces drawing error relative to the ideal polygon but also means that a filled polygon won&#146;t match the same polygon drawn unfilled. Narrow wedges and one-pixel-wide polygons will show up spottily. All in all, the choice of polygon-filling approach depends entirely on the ways in which the filled polygons must be used.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/i.jpg"><TD WIDTH="95%" VALIGN="TOP" ALIGN="LEFT"><SMALL><I>This sort of non-overlapping polygon filling isn&#146;t ideal for all purposes. Polygons are skewed toward the top and left edges, which not only introduces drawing error relative to the ideal polygon but also means that a filled polygon won&#146;t match the same polygon drawn unfilled. Narrow wedges and one-pixel-wide polygons will show up spottily. All in all, the choice of polygon-filling approach depends entirely on the ways in which the filled polygons must be used.</I></SMALL>
</TABLE>
<P>For our purposes, nonoverlapping polygons are the way to go, so let&#146;s have at them.
</P>

View file

@ -153,7 +153,7 @@ void DrawHorizontalLineList(struct HLineList * HLineListPtr,
<P>There&#146;s no secret as to why last chapter&#146;s <B>ScanEdge</B> was so slow: It used floating point calculations. One secret of fast graphics is using integer or fixed-point calculations, instead. (Sure, the floating point code would run faster if a math coprocessor were installed, but it would still be slower than the alternatives; besides, why require a math coprocessor when you don&#146;t have to?) Both integer and fixed-point calculations are fast. In many cases, fixed-point is faster, but integer calculations have one tremendous virtue: They&#146;re completely accurate. The tiny imprecision inherent in either fixed or floating-point calculations can result in occasional pixels being one position off from their proper location. This is no great tragedy, but after going to so much trouble to ensure that polygons don&#146;t overlap at common edges, why not get it exactly right?</P>
<P>In fact, when I tested out the integer edge tracing code by comparing an integer-based test image to one produced by floating-point calculations, two pixels out of the whole screen differed, leading me to suspect a bug in the integer code. It turned out, however, that&#146;s in those two cases, the floating point results were sufficiently imprecise to creep from just under an integer value to just over it, so that the <B>ceil</B> function returned a coordinate that was one too large.</P>
<TABLE WIDTH="100%"><TR>
<TD WIDTH="5%" ALIGN="LEFT" VALIGN="TOP"><IMG SRC="images/39-01i.jpg"><TD WIDTH="95%" VALIGN="TOP" ALIGN="LEFT"><SMALL><I>Floating point is very accurate&#151;but it is not precise. Integer calculations, properly performed, are.</I></SMALL>
<TD WIDTH="5%" ALIGN="LEFT" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%" VALIGN="TOP" ALIGN="LEFT"><SMALL><I>Floating point is very accurate&#151;but it is not precise. Integer calculations, properly performed, are.</I></SMALL>
</TABLE>
<P><BR></P>
<CENTER>

View file

@ -38,7 +38,7 @@
<P><BR></P>
<P>And indeed it does. When the test program is modified to draw to a local buffer, both the C and assembly language versions get 0.29 seconds faster, that being a measure of the time taken by display memory wait states. With those wait states factored out, the assembly language version of <B>DrawHorizontalLineList</B> becomes almost three times as fast as the C code.</P>
<TABLE WIDTH="100%"><TR>
<TD WIDTH="5%" ALIGN="LEFT" VALIGN="TOP"><IMG SRC="images/39-02i.jpg"><TD WIDTH="90%" VALIGN="TOP" ALIGN="LEFT"><SMALL><I>There is a lesson here. An optimization has no fixed payoff; its value fluctuates according to the context in which it is used. There&#146;s relatively little benefit to further optimizing code that already spends half its time waiting for display memory; no matter how good your optimizations, you&#146;ll get only a two-times speedup at best, and generally much less than that. There is, on the other hand, potential for tremendous improvement when drawing to system memory, so if that&#146;s where most of your drawing will occur, optimizations such as Listing 39.3 are well worth the effort.</I></SMALL>
<TD WIDTH="5%" ALIGN="LEFT" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="90%" VALIGN="TOP" ALIGN="LEFT"><SMALL><I>There is a lesson here. An optimization has no fixed payoff; its value fluctuates according to the context in which it is used. There&#146;s relatively little benefit to further optimizing code that already spends half its time waiting for display memory; no matter how good your optimizations, you&#146;ll get only a two-times speedup at best, and generally much less than that. There is, on the other hand, potential for tremendous improvement when drawing to system memory, so if that&#146;s where most of your drawing will occur, optimizations such as Listing 39.3 are well worth the effort.</I></SMALL>
<TR>
<TD VALIGN="TOP" ALIGN="LEFT">
<TD VALIGN="TOP" ALIGN="LEFT"><SMALL><I>Know the environments in which your code will run, and know where the cycles go in those environments.</I></SMALL>

View file

@ -38,7 +38,7 @@
<P><BR></P>
<P>An insertion sort that scans backward through the AET from the current edge rather than forward from the start of the AET could be quite a bit faster, because edges rarely move more than one or two positions through the AET. However, scanning backward requires a doubly linked list, rather than the singly linked list used in Listing 40.1. I&#146;ve chosen to use a singly linked list partly to minimize memory requirements (double-linking requires an extra pointer field) and partly because supporting back links would complicate the code a good bit. The main reason, though, is that the potential rewards for the complications of back links and insertion sorting aren&#146;t great enough; profiling a variety of polygons reveals that less than ten percent of total time is spent sorting the AET.
</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/40-01i.jpg"><TD WIDTH="95%"><SMALL><I>The potential 1 to 5 percent speedup gained by optimizing AET sorting just isn&#146;t worth it in any but the most demanding application&#151;a good example of the need to keep an overall perspective when comparing the theoretical characteristics of various approaches.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>The potential 1 to 5 percent speedup gained by optimizing AET sorting just isn&#146;t worth it in any but the most demanding application&#151;a good example of the need to keep an overall perspective when comparing the theoretical characteristics of various approaches.</I></SMALL>
</TABLE>
<H3><A NAME="Heading8"></A><FONT COLOR="#000077">Nonconvex Polygons</FONT></H3>
<P>Nonconvex polygons can be filled somewhat faster than complex polygons. Because edges never cross or switch positions with other edges once they&#146;re in the AET, the AET for a nonconvex polygon needs to be sorted only when new edges are added. In order for this to work, though, edges must be added to the AET in strict left-to-right order. Complications arise when dealing with two edges that start at the same point, because slopes must be compared to determine which edge is leftmost. This is certainly doable, but because of space limitations and limited performance returns, I haven&#146;t implemented this in Listing 40.1.

View file

@ -42,7 +42,7 @@
<P>After I wrote the columns on polygons in <I>Dr. Dobb&#146;s Journal</I> that became Chapters 38&#150;40, long-time reader Bill Huber wrote to take me to task&#151;and a well-deserved kick in the fanny it was, I might add&#151;for my use of non-standard polygon terminology in those columns. Unix&#146;s X-Window System (XWS) defines three categories of polygons: complex, nonconvex, and convex. These three categories, each a specialized subset of the preceding category, not-so-coincidentally map quite nicely to three increasingly fast polygon filling techniques. Therefore, I used the XWS names to describe the sorts of polygons that can be drawn with each of the polygon filling techniques.</P>
<P>The problem is that those names don&#146;t accurately describe all the sorts of polygons that the techniques are capable of drawing. Convex polygons are those for which no interior angle is greater than 180 degrees. The &#147;convex&#148; drawing approach described in the previous few chapters actually handles a number of polygons that are not convex; in fact, it can draw any polygon through which no horizontal line can be drawn that intersects the boundary more than twice. (In other words, the boundary reverses the Y direction exactly twice, disregarding polygons that have degenerated into horizontal lines, which I&#146;m going to ignore.)</P>
<P>Bill was kind enough to send me the pages out of <I>Computational Geometry, An Introduction</I> (Springer-Verlag, 1988) that describe the correct terminology; such polygons are, in fact, &#147;monotone with respect to a vertical line&#148; (which unfortunately makes a rather long <B>#define</B> variable). Actually, to be a tad more precise, I&#146;d call them &#147;monotone with respect to a vertical line and simple,&#148; where &#147;simple&#148; means &#147;not self-intersecting.&#148; Similarly, the polygon type I called &#147;nonconvex&#148; is actually &#147;simple,&#148; and I suppose what I called &#147;complex&#148; should be referred to as &#147;nonsimple,&#148; or maybe just &#147;none of the above.&#148;</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/41-01i.jpg"><TD WIDTH="95%"><SMALL><I>This may seem like nit-picking, but actually, it isn&#146;t; what it&#146;s really about is the tremendous importance of having a shared language. In one of his books, Richard Feynman describes having developed his own mathematical framework, complete with his own notation and terminology, in high school. When he got to college and started working with other people who were at his level, he suddenly understood that people can&#146;t share ideas effectively unless they speak the same language; otherwise, they waste a great deal of time on misunderstandings and explanation.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>This may seem like nit-picking, but actually, it isn&#146;t; what it&#146;s really about is the tremendous importance of having a shared language. In one of his books, Richard Feynman describes having developed his own mathematical framework, complete with his own notation and terminology, in high school. When he got to college and started working with other people who were at his level, he suddenly understood that people can&#146;t share ideas effectively unless they speak the same language; otherwise, they waste a great deal of time on misunderstandings and explanation.</I></SMALL>
</TABLE>
<P>Or, as Bill Huber put it, &#147;You are free to adopt your own terminology when it suits your purposes well. But you risk losing or confusing those who could be among your most astute readers&#151;those who already have been trained in the same or a related field.&#148; Ditto. Likewise. <I>D&#146;accord</I>. And <I>mea culpa</I> ; I shall endeavor to watch my language in the future.</P>
<H3><A NAME="Heading3"></A><FONT COLOR="#000077">Nomenclature in Action</FONT></H3>

View file

@ -312,7 +312,7 @@ _DrawWuLine endp
<!-- END CODE //-->
<H4 ALIGN="LEFT"><A NAME="Heading6"></A><FONT COLOR="#000077">Notes on Wu Antialiasing</FONT></H4>
<P>Wu antialiasing can be applied to any curve for which it&#146;s possible to calculate at each step the positions and intensities of two bracketing pixels, although the implementation will generally be nowhere near as efficient as it is for lines. However, Wu&#146;s article in <I>Computer Graphics</I> does describe an efficient algorithm for drawing antialiased circles. Wu also describes a technique for antialiasing solids, such as filled circles and polygons. Wu&#146;s approach biases the edges of filled objects outward. Although this is no good for adjacent polygons of the sort used in rendering, it&#146;s certainly possible to design a more accurate polygon-antialiasing approach around Wu&#146;s basic weighting technique. The results would not be quite so good as more sophisticated antialiasing techniques, but they would be much faster.</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/42-01i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>In general, the results obtained by Wu antialiasing are only so-so, by theoretical measures. Wu antialiasing amounts to a simple box filter placed over a fixed-point step approximation of a line, and that process introduces a good deal of deviation from the ideal. On the other hand, Wu notes that even a 10 percent error in intensity doesn&#146;t lead to noticeable loss of image quality, and for Wu-antialiased lines up to 1K pixels in length, the error is under 10 percent. If it looks good, it is good&#151;and it looks good.</I></SMALL>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>In general, the results obtained by Wu antialiasing are only so-so, by theoretical measures. Wu antialiasing amounts to a simple box filter placed over a fixed-point step approximation of a line, and that process introduces a good deal of deviation from the ideal. On the other hand, Wu notes that even a 10 percent error in intensity doesn&#146;t lead to noticeable loss of image quality, and for Wu-antialiased lines up to 1K pixels in length, the error is under 10 percent. If it looks good, it is good&#151;and it looks good.</I></SMALL>
</TABLE>
<P>With a 16-bit error accumulator, fixed-point inaccuracy becomes a problem for Wu-antialiased lines longer than 1K. For such lines, you should switch to using 32-bit error values, which would let you handle lines of any practical length.
</P>

View file

@ -75,7 +75,7 @@ mov byte ptr es:[di],0ffh
<!-- END CODE SNIP //-->
<TABLE WIDTH="100%">
<TR>
<TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/44-01i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><I><SMALL>If you&#146;re familiar with VGA programming, you&#146;re no doubt aware that everything that can be done with write mode 3 can also be accomplished in write mode 0 or write mode 2 by using the Bit Mask register. However, setting the Bit Mask register requires at least one <B>OUT</B> per byte written, in addition to the read and write of display memory, and <B>OUT</B>s are often slower than display memory accesses, especially on 386s and 486s. One of the great virtues of write mode 3 is that it requires virtually no <B>OUT</B>s and is therefore substantially faster for masking than the other write modes.</SMALL></I>
<TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><I><SMALL>If you&#146;re familiar with VGA programming, you&#146;re no doubt aware that everything that can be done with write mode 3 can also be accomplished in write mode 0 or write mode 2 by using the Bit Mask register. However, setting the Bit Mask register requires at least one <B>OUT</B> per byte written, in addition to the read and write of display memory, and <B>OUT</B>s are often slower than display memory accesses, especially on 386s and 486s. One of the great virtues of write mode 3 is that it requires virtually no <B>OUT</B>s and is therefore substantially faster for masking than the other write modes.</SMALL></I>
</TABLE>
<P>In short, write mode 3 is a good choice for single-color drawing that modifies individual pixels within display memory bytes. Not coincidentally, the sample application draws only single-color objects within the animation area; this allows write mode 3 to be used for all drawing, in keeping with our desire for speedy screen updates.
</P>

View file

@ -38,7 +38,7 @@
<P><BR></P>
<P>Given the above assumptions, drawing text is easy; we simply copy each byte of each character to the appropriate location in display memory, and <I>voila</I>, we&#146;re done. Text copying is done in write mode 0, in which the byte written to display memory is copied to all four planes at once; hence, 1-bits turn into white (color value 0FH, with 1-bits in all four planes), and 0-bits turn into black (color value 0). This is faster than using write mode 3 because write mode 3 requires a read/write of display memory (or at least preloading the latches with the background color), while the write mode 0 approach requires only a write to display memory.</P>
<TABLE WIDTH="100%"><TR>
<TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/44-02i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>Is write mode 0 always the best way to do text? Not at all. The write mode 0 approach described above draws both foreground and background pixels within the character box, forcing the background pixels to black at the same time that it forces the foreground pixels to white. If you want to draw transparent text (that is, draw only the character pixels, not the surrounding background box), write mode 3 is ideal. Also, matters get far more complicated if characters that aren&#146;t 8 pixels wide are drawn, or if characters are drawn starting at arbitrary pixel locations, without the multiple-of-8 column restriction, so that rotation and masking are required. Lastly, the Map Mask register can be used to draw text in colors other than white&#151;but only if the background is black. Otherwise, the data remaining in the planes protected by the Map Mask will remain and can interfere with the colors of the text being drawn.</I></SMALL>
<TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>Is write mode 0 always the best way to do text? Not at all. The write mode 0 approach described above draws both foreground and background pixels within the character box, forcing the background pixels to black at the same time that it forces the foreground pixels to white. If you want to draw transparent text (that is, draw only the character pixels, not the surrounding background box), write mode 3 is ideal. Also, matters get far more complicated if characters that aren&#146;t 8 pixels wide are drawn, or if characters are drawn starting at arbitrary pixel locations, without the multiple-of-8 column restriction, so that rotation and masking are required. Lastly, the Map Mask register can be used to draw text in colors other than white&#151;but only if the background is black. Otherwise, the data remaining in the planes protected by the Map Mask will remain and can interfere with the colors of the text being drawn.</I></SMALL>
</TABLE>
<P>I&#146;m not going to delve any deeper into the considerable issues of drawing VGA text; I just want to sensitize you to the existence of approaches other than the ones used in Listings 44.1 and 44.2. On the VGA, the rule is: If there&#146;s something you want to do, there probably are 10 ways to do it, each with unique strengths and weaknesses. Your mission, should you decide to accept it, is to figure out which one is best for your particular application.
</P>

View file

@ -85,7 +85,7 @@
<P>
</P>
<TABLE WIDTH="100%"><TR>
<TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/45-01i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I><B>OUT</B>s, in general, are lousy on the 486 (and to think they only took three cycles on the 286!). <B>OUT</B>s to VGAs are particularly lousy. Display memory performance is pretty poor, especially for reads. The conclusions are obvious, I would hope. Structure your graphics code, and, in general, all 486 code, to avoid <B>OUT</B>s.</I></SMALL>
<TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I><B>OUT</B>s, in general, are lousy on the 486 (and to think they only took three cycles on the 286!). <B>OUT</B>s to VGAs are particularly lousy. Display memory performance is pretty poor, especially for reads. The conclusions are obvious, I would hope. Structure your graphics code, and, in general, all 486 code, to avoid <B>OUT</B>s.</I></SMALL>
</TABLE>
<P>For graphics, this especially means using write mode 3 rather than the bit-mask register. When you must use the bit mask, arrange drawing so that you can set the bit mask once, then do a lot of drawing with that mask. For example, draw a whole edge at once, then the middle, then the other edge, rather than setting the bit mask several times on each scan line to draw the edge and middle bytes together. Don&#146;t read from display memory if you don&#146;t have to. Write each pixel once and only once.
</P>

View file

@ -38,7 +38,7 @@
<P><BR></P>
<P>One point I&#146;d like to make is that although the system-memory buffer in Listing 45.1 has exactly the same dimensions as the screen bitmap, that&#146;s not a requirement, and there are some good reasons not to make the two the same size. For example, if the system buffer is bigger than the area displayed on the screen, it&#146;s possible to pan the visible area around the system buffer. Or, alternatively, the system buffer can be just the size of a desired window, representing a window into a larger, virtual buffer. We could then draw the desired portion of the virtual bitmap into the system-memory buffer, then copy the buffer to the screen, and the effect will be of having panned the window to the new location.
</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/45-02i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>Another argument in favor of a small viewing window is that it restricts the amount of display memory actually drawn to. Restricting the display memory used for animation reduces the total number of display-memory accesses, which in turn boosts overall performance; it also improves the performance and appearance of panning, in which the whole window has to be redrawn or copied.</I></SMALL>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>Another argument in favor of a small viewing window is that it restricts the amount of display memory actually drawn to. Restricting the display memory used for animation reduces the total number of display-memory accesses, which in turn boosts overall performance; it also improves the performance and appearance of panning, in which the whole window has to be redrawn or copied.</I></SMALL>
</TABLE>
<P>If you keep a close watch, you&#146;ll notice that many high-performance animation games similarly restrict their full-featured animation area to a relatively small region. Often, it&#146;s hard to tell that this is the case, because the animation region is surrounded by flashy digitized graphics and by items such as scoreboards and status screens, but look closely and see if the animation region in your favorite game isn&#146;t smaller than you thought.
</P>

View file

@ -37,7 +37,7 @@
</CENTER>
<P><BR></P>
<TABLE WIDTH="100%">
<TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/47-01i.jpg"><TD WIDTH="95%" VALIGN="TOP"><SMALL><I>In general, performing plane-at-a-time operations can make almost any Mode X operation, at the worst, nearly as fast as the same operation in mode 13H (although this sort of Mode X programming is admittedly fairly complex). In this pursuit, it can help to organize data structures with Mode X in mind. For example, icons could be prearranged in system memory with the pixels organized into four plane-oriented sets (or, again, in four sets per scan line to avoid a fading-in effect) to facilitate copying to the screen a plane at a time with <B>REP MOVS</B></SMALL></I>
<TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%" VALIGN="TOP"><SMALL><I>In general, performing plane-at-a-time operations can make almost any Mode X operation, at the worst, nearly as fast as the same operation in mode 13H (although this sort of Mode X programming is admittedly fairly complex). In this pursuit, it can help to organize data structures with Mode X in mind. For example, icons could be prearranged in system memory with the pixels organized into four plane-oriented sets (or, again, in four sets per scan line to avoid a fading-in effect) to facilitate copying to the screen a plane at a time with <B>REP MOVS</B></SMALL></I>
</TABLE>
<P><B>LISTING 47.5 L47-5.ASM</B></P>
<!-- CODE //-->

View file

@ -52,7 +52,7 @@
<H3><A NAME="Heading4"></A><FONT COLOR="#000077">Copying Pixel Blocks within Display Memory</FONT></H3>
<P>Another fine use for the latches is copying pixels from one place in display memory to another. Whenever both the source and the destination share the same nibble alignment (that is, their start addresses modulo four are the same), it is not only possible but quite easy to use the latches to copy four pixels at a time. Listing 48.3 shows a routine that copies via the latches. (When the source and destination do not share the same nibble alignment, the latches cannot be used because the source and destination planes for any given pixel differ. In that case, you can set the Read Map register to select a source plane and the Map Mask register to select the corresponding destination plane. Then, copy all pixels in that plane, repeating for all four planes.)
</P>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/48-01i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Although copying through the latches is, in general, a speedy technique, especially on slower VGAs, it&#146;s not always a win. Reading video memory tends to be quite a bit slower than writing, and on a fast VLB or PCI adapter, it can be faster to copy from main memory to display memory than it is to copy from display memory to display memory via the latches.</I></SMALL>
<TABLE WIDTH="100%"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="5%"><IMG SRC="images/i.jpg"><TD ALIGN="LEFT" VALIGN="TOP" WIDTH="95%"><SMALL><I>Although copying through the latches is, in general, a speedy technique, especially on slower VGAs, it&#146;s not always a win. Reading video memory tends to be quite a bit slower than writing, and on a fast VLB or PCI adapter, it can be faster to copy from main memory to display memory than it is to copy from display memory to display memory via the latches.</I></SMALL>
</TABLE>
<P><BR></P>
<CENTER>

View file

@ -134,7 +134,7 @@ typedef struct {
<P>Listings 49.1 and 49.2, like all Mode X code I&#146;ve presented, perform no clipping, because clipping code would complicate the listings too much. While clipping can be implemented directly in the low-level Mode X routines (at the beginning of Listing 49.1, for instance), another, potentially simpler approach would be to perform clipping at a higher level, modifying the coordinates and dimensions passed to low-level routines such as Listings 49.1 and 49.2 as necessary to accomplish the desired clipping. It is for precisely this reason that the low-level Mode X routines support programmable start coordinates in the source images, rather than assuming (0,0); likewise for the distinction between the width of the image and the width of the area of the image to draw.
</P>
<P>Also, it would be more efficient to make up structures that describe the source and destination bitmaps, with dimensions and coordinates built in, and simply pass pointers to these structures to the low level, rather than passing many separate parameters, as is now the case. I&#146;ve used separate parameters for simplicity and flexibility.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/49-01i.jpg"><TD WIDTH="95%"><SMALL><I>Be aware that as nifty as Mode X hardware-assisted masked copying is, whether or not it&#146;s actually faster than software-only masked or transparent copying depends upon the processor and the video adapter. The advantage of Mode X masked copying is the 32-bit parallelism; the disadvantages are the need to read display memory and the need to perform an <B>OUT</B> for every four pixels. (<B>OUT</B> is a slow 486/Pentium instruction, and most VGAs respond to <B>OUT</B>s much more slowly than to display memory writes.)</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>Be aware that as nifty as Mode X hardware-assisted masked copying is, whether or not it&#146;s actually faster than software-only masked or transparent copying depends upon the processor and the video adapter. The advantage of Mode X masked copying is the 32-bit parallelism; the disadvantages are the need to read display memory and the need to perform an <B>OUT</B> for every four pixels. (<B>OUT</B> is a slow 486/Pentium instruction, and most VGAs respond to <B>OUT</B>s much more slowly than to display memory writes.)</I></SMALL>
</TABLE>
<P><BR></P>
<CENTER>

View file

@ -41,7 +41,7 @@
<H3><A NAME="Heading2"></A><FONT COLOR="#000077">Using Backface Removal to Eliminate Hidden Surfaces</FONT></H3>
<P>As I&#146;m fond of pointing out, computer animation isn&#146;t a matter of mathematically exact modeling or raw technical prowess, but rather of fooling the eye and the mind. That&#146;s especially true for 3-D animation, where we&#146;re not only trying to convince viewers that they&#146;re seeing objects on a screen&#151;when in truth that screen contains no objects at all, only gaggles of pixels&#151;but we&#146;re also trying to create the illusion that the objects exist in three-space, possessing four dimensions (counting movement over time as a fourth dimension) of their own. To make this magic happen, we must provide cues for the eye not only to pick out boundaries, but also to detect depth, orientation, and motion. This involves perspective, shading, proper handling of hidden surfaces, and rapid and smooth screen updates; the whole deal is considerably more difficult to pull off on a PC than 2-D animation.
</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/51-01i.jpg"><TD WIDTH="95%" VALIGN="TOP"><SMALL><I>In some senses, however, 3-D animation is easier than 2-D. Because there&#146;s more going on in 3-D animation, the eye and brain tend to make more assumptions, and so are more apt to see what they expect to see, rather than what&#146;s actually there.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%" VALIGN="TOP"><SMALL><I>In some senses, however, 3-D animation is easier than 2-D. Because there&#146;s more going on in 3-D animation, the eye and brain tend to make more assumptions, and so are more apt to see what they expect to see, rather than what&#146;s actually there.</I></SMALL>
</TABLE>
<P>If you&#146;re piloting a (virtual) ship through a field of thousands of asteroids at high speed, you&#146;re unlikely to notice if the more distant asteroids occasionally seem to go right through each other, or if the topographic detail on the asteroids&#146; surfaces sometimes shifts about a bit. You&#146;ll be busy viewing the asteroids in their primary role, as objects to be navigated around, and the mere presence of topographic detail will suffice; without being aware of it, you&#146;ll fill in the blanks. Your mind will see the topography peripherally, recognize it for what it is supposed to be, and, unless the landscape does something really obtrusive such as vanishing altogether or suddenly shooting a spike miles into space, you will see what you expect to see: a bunch of nicely detailed asteroids tumbling around you.
</P>

View file

@ -39,7 +39,7 @@
<H3><A NAME="Heading6"></A><FONT COLOR="#000077">Fast Texture Mapping: An Implementation</FONT></H3>
<P>As you might expect, I&#146;ve implemented DDA texture mapping in X-Sharp, and the changes are reflected in the X-Sharp archive in this chapter&#146;s subdirectory on the listings disk. Listing 56.1 shows the new header file entries, and Listing 56.2 shows the actual texture-mapped polygon drawer. The set-pixel routine that Listing 56.2 calls is a slight modification of the Mode X set-pixel routine from Chapter 47. In addition, INITBALL.C has been modified to create three texture-mapped polygons and define the texture bitmaps, and modifications have been made to allow the user to flip the axis of rotation. You will of course need the complete X-Sharp library to see texture mapping in action, but Listings 56.1 and 56.2 are the actual texture mapping code in its entirety.
</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/56-01i.jpg"><TD WIDTH="95%" VALIGN="TOP" ALIGN="LEFT"><SMALL><I>Here&#146;s a major tip: DDA texture mapping looks best on fast-moving surfaces, where the eye doesn&#146;t have time to pick nits with the shearing and aliasing that&#146;s an inevi table by-product of such a crude approach. Compile DEMO1 from the X-Sharp archive in this chapter&#146;s subdirectory of the listings disk, and run it. The initial display looks okay, but certainly not great, because the rotational speed is so slow. Now press the S key a few times to speed up the rotation and flip between different rotation axes. I think you&#146;ll be amazed at how much better DDA texture mapping looks at high speed. This technique would be great for mapping textures onto hurtling asteroids or jets, but would come up short for slow, finely detailed movements.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/i.jpg"><TD WIDTH="95%" VALIGN="TOP" ALIGN="LEFT"><SMALL><I>Here&#146;s a major tip: DDA texture mapping looks best on fast-moving surfaces, where the eye doesn&#146;t have time to pick nits with the shearing and aliasing that&#146;s an inevi table by-product of such a crude approach. Compile DEMO1 from the X-Sharp archive in this chapter&#146;s subdirectory of the listings disk, and run it. The initial display looks okay, but certainly not great, because the rotational speed is so slow. Now press the S key a few times to speed up the rotation and flip between different rotation axes. I think you&#146;ll be amazed at how much better DDA texture mapping looks at high speed. This technique would be great for mapping textures onto hurtling asteroids or jets, but would come up short for slow, finely detailed movements.</I></SMALL>
</TABLE>
<P><B>LISTING 56.1 L56-1.C</B></P>
<!-- CODE //-->

View file

@ -47,7 +47,7 @@
</P>
<P>The obvious lesson here is that adequate speed is important to convincing animation. There&#146;s another, less obvious side to this lesson, though. I&#146;d been running the texture-mapping demo on a 20 MHz 386 with a slow VGA when I discovered the beneficial effects of greater animation speed. When, some time later, I ran the demo on a 33 MHz 486 with a fast VGA, I found that the faster rotation was too fast! The ball spun so rapidly that the eye couldn&#146;t blend successive images together into continuous motion, much like watching a badly flickering movie.</P>
<TABLE WIDTH="100%"><TR>
<TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/57-01i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>So the second lesson is that either too little or too much speed can destroy the illusion. Unless you&#146;re antialiasing, you need to tune the shifting of your images so that they&#146;re in the &#147;sweet spot&#148; of apparent motion, in which the eye is willing to ignore the jumping and aliasing, and blend the images together into continuous motion. Only experience can give you a feel for that sweet spot.</I></SMALL>
<TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>So the second lesson is that either too little or too much speed can destroy the illusion. Unless you&#146;re antialiasing, you need to tune the shifting of your images so that they&#146;re in the &#147;sweet spot&#148; of apparent motion, in which the eye is willing to ignore the jumping and aliasing, and blend the images together into continuous motion. Only experience can give you a feel for that sweet spot.</I></SMALL>
</TABLE>
<H4 ALIGN="LEFT"><A NAME="Heading4"></A><FONT COLOR="#000077">Fixed-Point Arithmetic, Redux</FONT></H4>
<P>In the previous chapter I added texture mapping to X-Sharp, but lacked space to explain some of its finer points. I&#146;ll pick up the thread now and cover some of those points here, and discuss the visual and performance enhancements that previous chapter&#146;s code needed&#151;and which are now present in the version of X-Sharp in this chapter&#146;s subdirectory on the CD-ROM.

View file

@ -38,7 +38,7 @@
<P><BR></P>
<P>I&#146;d like to emphasize that algorithmically and conceptually, there is <I>no</I> difference between scanning out a polygon top to bottom and scanning it out left to right; it is only in conjunction with the hardware organization of Mode X that the scanning direction matters in the least.</P>
<TABLE WIDTH="100%"><TR>
<TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/58-01i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>That&#146;s what Zen programming is all about, though; tying together two pieces of seemingly unrelated information to good effect&#151;and that&#146;s what I had failed to do. Like Robert Heinlein&#151;like all of us&#151;I had viewed the world through a filter composed of my ingrained assumptions, and one of those assumptions, based on all my past experience, was that pixel processing proceeds left to right. Eventually, I might have come up with Chris&#146;s approach; but I would only have come up with it when and if I relaxed and stepped back a little, and allowed myself&#151;almost dared myself&#151;to think of it. When you&#146;re optimizing, be sure to leave quiet, nondirected time in which to conjure up those less obvious solutions, and periodically try to figure out what assumptions you&#146;re making&#151;and then question them!</I></SMALL>
<TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>That&#146;s what Zen programming is all about, though; tying together two pieces of seemingly unrelated information to good effect&#151;and that&#146;s what I had failed to do. Like Robert Heinlein&#151;like all of us&#151;I had viewed the world through a filter composed of my ingrained assumptions, and one of those assumptions, based on all my past experience, was that pixel processing proceeds left to right. Eventually, I might have come up with Chris&#146;s approach; but I would only have come up with it when and if I relaxed and stepped back a little, and allowed myself&#151;almost dared myself&#151;to think of it. When you&#146;re optimizing, be sure to leave quiet, nondirected time in which to conjure up those less obvious solutions, and periodically try to figure out what assumptions you&#146;re making&#151;and then question them!</I></SMALL>
</TABLE>
<P><A NAME="Fig3"><!-- </A><A HREF="javascript:displayWindow('images/58-03.jpg',407,145 )"> --><IMG SRC="images/58-03.jpg"><BR><!-- </A>
<BR><A HREF="javascript:displayWindow('images/58-03.jpg',407,145)"> --><FONT COLOR="#000077"><B>Figure 58.3</B></FONT></A>&nbsp;&nbsp;<I>Texture mapping a single vertical column.</I>

View file

@ -97,7 +97,7 @@ void WalkBSPTree(NODE *pNode)
<!-- END CODE //-->
<TABLE WIDTH="100%">
<TR>
<TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/59-01i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>Be aware that BSP trees can often be made smaller and more efficient by detecting collinear surfaces (like aligned wall segments) and generating only one BSP node for each collinear set, with the collinear surfaces stored in, say, a linked list attached to that node. Collinear surfaces partition space identically and can&#146;t occlude one another, so it suffices to generate one splitting node for each collinear set.</I></SMALL>
<TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>Be aware that BSP trees can often be made smaller and more efficient by detecting collinear surfaces (like aligned wall segments) and generating only one BSP node for each collinear set, with the collinear surfaces stored in, say, a linked list attached to that node. Collinear surfaces partition space identically and can&#146;t occlude one another, so it suffices to generate one splitting node for each collinear set.</I></SMALL>
</TABLE>
<P><BR></P>
<CENTER>

View file

@ -40,7 +40,7 @@
</P>
<P>Let&#146;s look at it another way. All the code in Listing 59.2 does is say: &#147;Here I am at a node. First I&#146;ll visit the left subtree if there is one, then I&#146;ll visit this node, then I&#146;ll visit the right subtree if there is one. While I&#146;m visiting the left subtree, I&#146;ll just push a marker on a stack that tells me to come back here when the left subtree is done. If, after visiting a node, there are no right children to visit and nothing left on the stack, I&#146;m finished. The code does this at each node&#151;and that&#146;s <I>all</I> it does. That&#146;s all Listing 59.4 does, too, but people tend to get tangled up in pushes and pops and <B>while</B> loops when they use data recursion. When the implementation model changes to one with which they are unfamiliar, they abandon the perfectly good model they used before and try to rederive it in the new context by the seat of their pants.</P>
<TABLE WIDTH="100%"><TR>
<TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/59-02i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>Here&#146;s a secret when you&#146;re faced with a situation like this: Step back and get a clear picture of what your code has to do. Omit no steps. You should build a model that is so consistent and solid that you can instantly answer any question about how the code should behave in any situation. For example, my interviewees often decide, by trial and error, that there are two distinct types of right children: Right children visited after popping back to visit a node after the left subtree has been visited, and right children visited after descending to a node that has no left child. This makes the traversal code a mass of special cases, each of which has to be detected by the programmer by trying out scenarios. Worse, you can never be sure with this approach that you&#146;ve caught all the special cases.</I></SMALL>
<TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>Here&#146;s a secret when you&#146;re faced with a situation like this: Step back and get a clear picture of what your code has to do. Omit no steps. You should build a model that is so consistent and solid that you can instantly answer any question about how the code should behave in any situation. For example, my interviewees often decide, by trial and error, that there are two distinct types of right children: Right children visited after popping back to visit a node after the left subtree has been visited, and right children visited after descending to a node that has no left child. This makes the traversal code a mass of special cases, each of which has to be detected by the programmer by trying out scenarios. Worse, you can never be sure with this approach that you&#146;ve caught all the special cases.</I></SMALL>
<TR>
<TD>
<TD><SMALL><I>The alternative is to develop and apply a unifying model. There aren&#146;t really two types of right children; the rule is that all right children are visited after their parents are visited, period. The presence or absence of a left child is irrelevant. The possibility that a right child may be reached via different code paths depending on the presence of a left child does not affect the overall model. While this distinction may seem trivial it is in fact crucial, because if you have the model down cold, you can always tell if the implementation is correct by comparing it with the model.</I></SMALL>

View file

@ -48,7 +48,7 @@
</P>
<P ALIGN="RIGHT">(eq. 7)</P>
<P>(Note that the cross product operation is denoted by an X.) Unlike the dot product, the result of the cross product is a vector. Not just any vector, either; the vector generated by the cross product is perpendicular to both of the original vectors. Thus, the cross product can be used to generate a normal to any surface for which you have two vectors that lie within the surface. This means that we can generate the screenspace normals we need by taking the cross product of two adjacent polygon edges, as shown in Figure 61.4.</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/61-01i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>In fact, we can cull with only one-third the work needed to generate a full cross product; because we&#146;re interested only in the sign of the z component of the normal, we can skip entirely calculating the x and y components. The only caveat is to be careful that neither edge you choose is zero-length and that the edges aren&#146;t collinear, because the dot product can&#146;t produce a normal in those cases.</SMALL></I>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>In fact, we can cull with only one-third the work needed to generate a full cross product; because we&#146;re interested only in the sign of the z component of the normal, we can skip entirely calculating the x and y components. The only caveat is to be careful that neither edge you choose is zero-length and that the edges aren&#146;t collinear, because the dot product can&#146;t produce a normal in those cases.</SMALL></I>
</TABLE>
<P><A NAME="Fig4"><!-- </A><A HREF="javascript:displayWindow('images/61-04.jpg',407,283 )"> --><IMG SRC="images/61-04.jpg"><BR><!-- </A>
<BR><A HREF="javascript:displayWindow('images/61-04.jpg',407,283)"> --><FONT COLOR="#000077"><B>Figure 61.4</B></FONT></A>&nbsp;&nbsp;<I>How the cross product of polygon edge vectors generates a polygon normal.</I>

View file

@ -97,7 +97,7 @@ void LineIntersectPlane (float *linestart, float *lineend,
</P>
<P>Their approach is this: Think of rotation as projecting coordinates onto new axes. That is, given that you have points in, say, worldspace, define the new coordinate space (viewspace, for example) you want to rotate to by a set of three orthogonal unit vectors defining the new axes, and then project each point onto each of the three axes to get the coordinates in the new coordinate space, as shown for the 2-D case in Figure 61.8. In 3-D, this involves three dot products per point, one to project the point onto each axis. Translation can be done separately from rotation by simple addition.
</P>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/61-02i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>Rotation by projection is exactly the same as rotation via matrix multiplication; in fact, the rows of a rotation matrix <I>are</I> the orthogonal unit vectors pointing along the new axes. Rotation by projection buys us no technical advantages, so that&#146;s not what&#146;s important here; the key is that the concept of rotation by projection, together with a separate translation step, gives us a new way to look at transformation that I, for one, find easier to visualize and experiment with. A new frame of reference for how we think about 3-D frames of reference, if you will.</SMALL></I>
<TABLE WIDTH="100%"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="5%"><IMG SRC="images/i.jpg"><TD VALIGN="TOP" ALIGN="LEFT" WIDTH="95%"><SMALL><I>Rotation by projection is exactly the same as rotation via matrix multiplication; in fact, the rows of a rotation matrix <I>are</I> the orthogonal unit vectors pointing along the new axes. Rotation by projection buys us no technical advantages, so that&#146;s not what&#146;s important here; the key is that the concept of rotation by projection, together with a separate translation step, gives us a new way to look at transformation that I, for one, find easier to visualize and experiment with. A new frame of reference for how we think about 3-D frames of reference, if you will.</SMALL></I>
</TABLE>
<P>Three things I&#146;ve learned over the years are that it never hurts to learn a new way of looking at things, that it helps to have a clearer, more intuitive model in your head of whatever it is you&#146;re working on, and that new tools, or new ways to use old tools, are Good Things. My experience has been that rotation by projection, and dot product tricks in general, offer those sorts of benefits for 3-D.
</P>

View file

@ -46,7 +46,7 @@
<P>As it turns out, the <B>memcpy()</B> routine in the DOS version of our compiler (gcc) inexplicably copies memory a byte at a time. With my new routine, the non-page-flipped approach suddenly became slightly <I>faster</I> than page flipping.</P>
<P>The first relevant rule is pretty obvious: <I>Assume nothing</I>. Measure early and often. Know what&#146;s really going on when your program runs, if you catch my drift. To do otherwise is to risk looking mighty foolish.</P>
<P>The second rule: When you do look foolish (and trust me, it <I>will</I> happen if you do challenging work) have a good laugh at yourself, and use it as a reminder of Rule #1. I hadn&#146;t done any extra page-flipping work yet, so I didn&#146;t waste any time due to my faulty assumption that <B>memcpy()</B> performed a maximum-speed copy, but that was just luck. I should have done experiments until I was sure I knew what was going on before drawing any conclusions and acting on them.</P>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/62-01i.jpg"><TD WIDTH="95%" VALIGN="TOP" ALIGN="LEFT"><SMALL><I>In general, make it a point not to fall into a tightly focused rut; stay loose and think of alternative possibilities and new approaches, and always, always, always keep asking questions. It&#146;ll pay off big in the long run. If I hadn&#146;t indulged my curiosity by running the Pentium counter test on the copy to the screen, even though there was no specific reason to do so, I would never have discovered the <B>memcpy()</B> problem&#151;and by so doing I doubled the performance of the entire program in five minutes, a rare accomplishment indeed.</I></SMALL>
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP" ALIGN="LEFT"><IMG SRC="images/i.jpg"><TD WIDTH="95%" VALIGN="TOP" ALIGN="LEFT"><SMALL><I>In general, make it a point not to fall into a tightly focused rut; stay loose and think of alternative possibilities and new approaches, and always, always, always keep asking questions. It&#146;ll pay off big in the long run. If I hadn&#146;t indulged my curiosity by running the Pentium counter test on the copy to the screen, even though there was no specific reason to do so, I would never have discovered the <B>memcpy()</B> problem&#151;and by so doing I doubled the performance of the entire program in five minutes, a rare accomplishment indeed.</I></SMALL>
</TABLE>
<P>By the way, I have found the Pentium&#146;s performance counters to be very useful in of information on the performance counters and other aspects of the Pentium is Mike Schmit&#146;s book, <I>Pentium Processor Optimization Tools</I>, AP Professional, ISBN 0-12-627230-1.</P>
<P>Onward to rendering from a BSP tree.</P>

Some files were not shown because too many files have changed in this diff Show more