Replace invalid characters with HTML entities
— with — ’ with ’ + with + × with x ç with ç “ with “ ” with ” ‘ with ‘ • with • – with - µ with µ † with † Fix C++ θ with θ Yen symbol instead of times Fix broken apos Bullet again E-circumflex
This commit is contained in:
parent
9e52de9586
commit
500e7f5654
353 changed files with 5091 additions and 5091 deletions
10
02-01.html
10
02-01.html
|
|
@ -39,11 +39,11 @@
|
|||
<H2><A NAME="Heading1"></A><FONT COLOR="#000077">Chapter 2<BR>A World Apart
|
||||
</FONT></H2>
|
||||
<H3><A NAME="Heading2"></A><FONT COLOR="#000077">The Unique Nature of Assembly Language Optimization</FONT></H3>
|
||||
<P>As I showed in the previous chapter, optimization is by no means always a matter of “dropping into assembly.” In fact, in performance tuning high-level language code, assembly should be used rarely, and then only after you’ve made sure a badly chosen or clumsily implemented algorithm isn’t eating you alive. Certainly if you use assembly at all, make absolutely sure you use it <I>right</I>. The potential of assembly code to run <I>slowly</I> is poorly understood by a lot of people, but that potential is great, especially in the hands of the ignorant.</P>
|
||||
<P>Truly great optimization, however, happens <I>only</I> at the assembly level, and it happens in response to a set of dynamics that is totally different from that governing C/C<SMALL>++</SMALL> or Pascal optimization. I’ll be speaking of assembly-level optimization time and again in this book, but when I do, I think it will be helpful if you have a grasp of those assembly specific dynamics.</P>
|
||||
<P>As I showed in the previous chapter, optimization is by no means always a matter of “dropping into assembly.” In fact, in performance tuning high-level language code, assembly should be used rarely, and then only after you’ve made sure a badly chosen or clumsily implemented algorithm isn’t eating you alive. Certainly if you use assembly at all, make absolutely sure you use it <I>right</I>. The potential of assembly code to run <I>slowly</I> is poorly understood by a lot of people, but that potential is great, especially in the hands of the ignorant.</P>
|
||||
<P>Truly great optimization, however, happens <I>only</I> at the assembly level, and it happens in response to a set of dynamics that is totally different from that governing C/C<SMALL>++</SMALL> or Pascal optimization. I’ll be speaking of assembly-level optimization time and again in this book, but when I do, I think it will be helpful if you have a grasp of those assembly specific dynamics.</P>
|
||||
<P>As usual, the best way to wade in is to present a real-world example.</P>
|
||||
<H3><A NAME="Heading3"></A><FONT COLOR="#000077">Instructions: The Individual versus the Collective</FONT></H3>
|
||||
<P>Some time ago, I was asked to work over a critical assembly subroutine in order to make it run as fast as possible. The task of the subroutine was to construct a nibble out of four bits read from different bytes, rotating and combining the bits so that they ultimately ended up neatly aligned in bits 3-0 of a single byte. (In case you’re curious, the object was to construct a 16-color pixel from bits scattered over 4 bytes.) I examined the subroutine line by line, saving a cycle here and a cycle there, until the code truly seemed to be optimized. When I was done, the key part of the code looked something like this:
|
||||
<P>Some time ago, I was asked to work over a critical assembly subroutine in order to make it run as fast as possible. The task of the subroutine was to construct a nibble out of four bits read from different bytes, rotating and combining the bits so that they ultimately ended up neatly aligned in bits 3-0 of a single byte. (In case you’re curious, the object was to construct a 16-color pixel from bits scattered over 4 bytes.) I examined the subroutine line by line, saving a cycle here and a cycle there, until the code truly seemed to be optimized. When I was done, the key part of the code looked something like this:
|
||||
</P>
|
||||
<!-- CODE SNIP //-->
|
||||
<PRE>
|
||||
|
|
@ -57,7 +57,7 @@ LoopTop:
|
|||
jnz LoopTop ;process the next bit, if any
|
||||
</PRE>
|
||||
<!-- END CODE SNIP //-->
|
||||
<P>Now, it’s hard to write code that’s much faster than seven instructions, only one of which accesses memory, and most programmers would have called it a day at this point. Still, something bothered me, so I spent a bit of time going over the code again. Suddenly, the answer struck me—the code was rotating each bit into place separately, so that a multibit rotation was being performed every time through the loop, for a total of four separate time-consuming multibit rotations!
|
||||
<P>Now, it’s hard to write code that’s much faster than seven instructions, only one of which accesses memory, and most programmers would have called it a day at this point. Still, something bothered me, so I spent a bit of time going over the code again. Suddenly, the answer struck me—the code was rotating each bit into place separately, so that a multibit rotation was being performed every time through the loop, for a total of four separate time-consuming multibit rotations!
|
||||
</P>
|
||||
<TABLE WIDTH="100%"><TD WIDTH="5%" VALIGN="TOP"><IMG SRC="images/i.jpg"><TD WIDTH="95%"><SMALL><I>While the instructions themselves were individually optimized, the overall approach did not make the best possible use of the instructions.</I></SMALL>
|
||||
</TABLE>
|
||||
|
|
@ -76,7 +76,7 @@ LoopTop:
|
|||
; positions at the same time
|
||||
</PRE>
|
||||
<!-- END CODE //-->
|
||||
<P>This moved the costly multibit rotation out of the loop so that it was performed just once, rather than four times. While the code may not look much different from the original, and in fact still contains exactly the same number of instructions, the performance of the entire subroutine improved by about 10 percent from just this one change. (Incidentally, that wasn’t the end of the optimization; I eliminated the <B>DEC</B> and <B>JNJ</B> instructions by expanding the four iterations of the loop—but that’s a tale for another chapter.)</P>
|
||||
<P>This moved the costly multibit rotation out of the loop so that it was performed just once, rather than four times. While the code may not look much different from the original, and in fact still contains exactly the same number of instructions, the performance of the entire subroutine improved by about 10 percent from just this one change. (Incidentally, that wasn’t the end of the optimization; I eliminated the <B>DEC</B> and <B>JNJ</B> instructions by expanding the four iterations of the loop—but that’s a tale for another chapter.)</P>
|
||||
<P>The point is this: To write truly superior assembly programs, you need to know what the various instructions do and which instructions execute fastest...and more. You must also learn to look at your programming problems from a variety of perspectives so that you can put those fast instructions to work in the most effective ways.</P>
|
||||
<H3><A NAME="Heading4"></A><FONT COLOR="#000077">Assembly Is Fundamentally Different</FONT></H3>
|
||||
<P>Is it really so hard as all that to write good assembly code for the PC? Yes! Thanks to the decidedly quirky nature of the x86 family CPUs, assembly language differs fundamentally from other languages, and is undeniably harder to work with. On the other hand, the potential of assembly code is much greater than that of other languages, as well.
|
||||
|
|
|
|||
Loading…
Reference in a new issue