Showing posts with label Memory Management. Show all posts
Showing posts with label Memory Management. Show all posts

Monday, January 26, 2009

Programming Techniques for Memory Management (Part 2) by Roger Donnay

This section written by Roger Donnay is the second part of CA-Clipper 5.x Memory Management by Roger Donnay & Jud Cole.


The Affect of Databases on Memory Usage

Many programmers have taken an affection to "data-driven" programming to speed up the development of custom applications. Since CA-Clipper evolved as an X-Base type language, the common approach to data-driven programming is to create an "engine" or "kernel" .EXE program that retrieves "custom" information about the application from a set of database files. The more sophisticated the system, the more data files are required for configuring the application. It is important to understand the impact of databases on memory usage when designing applications that use many databases.

When a database is opened with the USE command, all the field names are placed into the public symbol table. This allows database field names to be used in expressions in the same manner as memvars. Because this symbol table is PUBLIC, the field names will remain in the symbol table, even after the database is closed. CA-Clipper employs no mechanism to remove symbols from the symbol table, only to add them. As each database is opened the symbols are added, thereby reducing the remaining available root memory ( MEMORY(0) ). It is conceivable that an application could run out of memory if many databases are opened and closed. Fortunately, symbols of the same name are "reused" in the symbol table, so if you open and close the same database many times, the symbol memory space will not be reduced. Keeping this in mind, it is a good practice to use databases with fields of the same name.

To improve speed operation, some programmers will open data-dictionary databases at the start-up of the program, load the custom program configuration into arrays, then close the databases. This action will reduce the amount of memory needed for data structures and file buffers and lower the number of DOS handles required, but the symbols will still be added to the symbol table even though they may never be accessed by the program. A data-driven application in which the databases are used only during the start-up of the application could be re-designed to convert the database information to a text file to generate a "runtime" version of the application that will load the array(s) from a text file rather than databases, thereby eliminating the symbol-table problem.

Declare FIELDS in your Compiled Codes

To prevent the need to add Field Names to the symbol table at runtime, it is always a good idea to declare fields to the Compiler by using the FIELD statement in your source code. This will insure that the symbol table memory is allocated at the time the application is linked into an executable program rather than during the running of the program.

Managing Code Segment Size ( C/ASM)

If you are writing code in C or Assembly, then the linker's overlay manager must bring an entire code segment into memory at one time. If you have many large code segments, then the size of the overlay pool may need to be increased. Normally, the requirements of the application dictate the size of the code segments, but it should be noted that a little extra time in optimizing the size of your code will usually pay off if you have a large library of routines.

CA-Clipper-compiled P-Code is loaded in fixed-size pages so this is not a requirement when programming in CA-Clipper.

Reducing RDD Memory

Try to use as few database drivers in your application as is necessary. For example, if you choose to use the DBFCDX driver for compatability with with FoxPro files, then make sure that your application does not link in the DBFNTX (default) driver unless you must have .DBT or .NTX compatability. This is accomplished by modifying the source code in the RDDSYS.PRG that is included with CA-Clipper.

Using the "Garbage Collector"

A variety of memory problems can be caused by memory "fragmentation". Fragmentation can be caused when symbols are added to the symbol table or other "locked" memory is allocated before other "unlocked" memory segments can be released. Database languages require that the symbol table be non-contiguous and not pre-allocated but instead must be allowed to "grow" during the application. This "late binding" is the essence and power behind data-driven engines (like many Clipper applications), unfortunately this kind of flexibility can cause symbols and work-area structures to be added haphazardly thruout the conventional memory space. MEMORY(0) can report that there is plenty of memory available but after the program runs for while it exists in many small segments rather that the single large segment that was available at the start of the application.

When a large segment of memory is required by the application and is not available, Clipper has to try to move around segments that are not "locked" so it can get sufficient memory for the routine that's being called. This is "garbage collection". Unfortunately, symbol memory and other memory allocation are "locked" and cannot be moved. Applications should take into consideration reducing the amount of fixed heap, stack, and VMM memory that gets allocated during the running of the program, and also to insure that garbage collection is forced on a methodical basis to prevent the fragmention.

CA-Clipper uses a technique called "scavenging" in its memory garbage collector. The garbage collector is automatically invoked during writes to the screen and "wait" states while waiting for keyboard input. Programs that have routines which have few screen-write or input routines can fragment memory badly particularly if a lot of file opening and closing, reindexing, appending, etc. is going on with little operator interaction. It is recommended that you put calls to Devout("") in your code to invoke the garbage collector in the event that you experience "Conventional Memory Exhausted", "Memory Low", or "Stack Eval" errors. My own experience with this method has not always given me desirable results, however, and I have found that placing calls to MEMORY(-1) in strategic locations in my code, such as just before opening databases, or creating large arrays actually produces better results. This method is not supported by CA or many other developers so use this recommendation with caution. Another popular method of perfoming garbage collection is to use the FT_IDLE() function from the public domain NANFORUM TOOLKIT available on many BBS's including the CLIPPER forum of COMPUSERVE. Garbage collection should be forced especially before opening databases and indexes.

Managing DGROUP Memory

There is an area of memory in CA-Clipper, as in all large model languages, referred to as DGROUP. This 64k region of low (root) memory is also called the DS (Data Segment) because it is pointed to by the DS and SS registers. Data is often stored in this area because it can be accessed faster than data stored elsewhere. DGROUP is essentially a NEAR CALL area because accessing information in this area is accomplished by a one-word address rather than the more common 2-word addressing required to access other data areas.

Assembly-language programmers like to store data in the DGROUP area to improve performance of their application. Data in this area doesn't get overlayed and it is insured that it will always be accessible regardless of the state of the application. Interrupt handlers always store data in DGROUP to insure that the data will be accessible in the event of an interrupt. CA-Clipper stores it's stacks and static values in DGROUP.

DGROUP is structured like this.

direction of growth | --> <-- |

+--------------+-----------+------------+---------+----------------+ | C/ASM static | cpu stack | eval stack | "space" | Clipper static | +--------------+-----------+------------+---------+----------------+

The eval stack is where local and private variable VALUEs are stored. The CA-Clipper static area is where CA-Clipper static VALUES are stored. ITEMs generated by programs using ITEM.API also create VALUE entries in the CA-Clipper static area. Each VALUE entry uses 14 bytes.

If the eval stack grows into the CA-Clipper static area (or vice versa), you get a UE 667 (stack fault) error.

In low memory situations, the VMM will allocate "space" to the conventional memory pool. After this happens, if the eval stack grows into the allocated space, you get a UE 668 error. Likewise, if the CA-Clipper static area needs to grow, you get a UE 669 error.

When the eval stack or CA-Clipper static area expands, and later retracts, the system maintains "watermarks" to indicate the farthest expansion of them. The VMM allocates only the space between the "watermarks". So one way to control things is to make sure the "watermarks" are set to allow your program to execute normally.

There is very little that a CA-Clipper programmer can do to resolve DGROUP problems other than to avoid using third-party products and/or C/ASM code that uses SMALL MODEL rather than LARGE MODEL programming techniques and replace LOCAL variables with LOCAL arrays. CA-Clipper is a LARGE MODEL programming language and it is recommended in the CA-Clipper API that extensions to the language should also be compiled as LARGE MODEL. Unfortunately, many C/ASM programmers develop libraries designed for speed performance rather than memory performance. I refer to these products as "DGROUP Hogs". These libraries will consume so much of the C/ASM static space that there will be literally nothing left to run the application. If you find that you cannot limit the usage of third-party libraries, then and alternative solution is to use one or more of the methods described below under "Managing the EVAL Stack".

There are several methods to determine how much DGROUP a library uses. See next page:

Monitor memory with //Info

When you start your CA-Clipper application with the //INFO option, CA-Clipper reports the condition of memory at the start of the application. The reports looks like this:

DS=4F4E:0000 DS avail=30KB OS avail=255KB EMM avail=960KB

DS= is the starting address of the DGROUP segment.

DS avail=KB reports the amount of DGROUP available.

OS avail=KB reports the amount of conventional memory available.

EMM avail=KB report the amount of expanded memory allocated to the current application.

When the DS avail (DGROUP) is less than 15k at the start-up of an application, some large applications could experience stack eval errors at runtime. This number is arbitrary and your application may actually run just fine with less available DGROUP, however it is important that you monitor DS especially if your application uses third-party libraries or lots of LOCAL and STATIC memvars. It's always a good idea to make note of the amount of DS avail in your current application before and after you add new functions or libraries to the application. If you suddenly notice a dramatic decrease in DS avail after adding new code, then you should take note that you may be using a function from a library that uses excessive DGROUP.

Create a .MAP file

Rtlink and Blinker support a MAP= command that creates a map file with information about segment usage. This map file will show you which routines use DGROUP and exactly how much they consume.

Managing the EVAL Stack

If your application uses up a lot of DGROUP memory, then you may be required to re-visit your code and either reduce the hit on the EVAL stack or expand the size of the EVAL stack. In my opinion, it is quite easy to resolve these problems with minor changes to source code.

Use LOCAL ARRAYS instead of LOCAL VARIABLES

Every time a new function is called, all LOCAL variables are "pushed" onto the EVAL stack. They are "popped" off the stack when returning to the calling program. If a procedure or function is called recursively, then it can be quite easy to blow up the eval stack with a 667 Stack Eval Error. For example, I had this problem in one of my larger applications due to the fact that programs were deeply nested via calls to a common set of menu functions. I found that my main menu function could blow up the stack if called recursively as few as 5 times. Each time this function was called, 85 LOCAL variables were pushed on to the eval stack. I completely eliminated the problem with about 2 hours of programming and debugging by replacing the 85 separate LOCAL variables with 1 LOCAL array. A LOCAL array uses the same amount of stack space as any other LOCAL variable, therefore I reduced the amount of values pushed on the stack from 85 to 1. After doing this, I could not blow up the stack even if I called the menu system recursively more than 30 times. This was accomplished quite easily with very minor code changes. Here is an example of the BEFORE and AFTER code.

BEFORE:

FUNCTION DC_MenuMain( )

LOCAL cColor, cMenuScrn, nInkey, cTempStr, nTempNum, nItemLen,; cOldColor, nItems, nMsgLen, nHotStart, cBoxColor, ; cMenuColor, cHotColor, cSelColor, cBuffer, nRow, nCol, ; lMouseHit, nSaveRow, nSaveCol, lMessage, cGrayColor, ; nMouseStat, nStRow, nStCol, nEnRow, nEnCol, lMouVisible, ; nPass, aSubMenu, aSubBlock, nElement, aMenuItems, ; aHotKeys, aMenuBlocks, aSubItems, aBlockItems, cType, ; nStart, cTitle, lBar, lReturnVal, lShadow, aMouseKeys, ....

AFTER:

FUNCTION DC_MenuMain( )

LOCAL aMenu := Array(85)

#define cColor aMenu[1] ; #define cMenuScrn aMenu[2]

#define nInkey aMenu[3] ; #define cTempStr aMenu[4]

#define nTempNum aMenu[5] ; #define nItemLen aMenu[6]

#define cOldColor aMenu[7] ; #define nItems aMenu[8]

#define nMsgLen aMenu[9] ; #define nItems aMenu[10]

#define nHotStart aMenu[11] ; #define cBoxColor aMenu[12]

#define cMenuColor aMenu[13] ; #define cHotColor aMenu[14]

#define cSelColor aMenu[15] ; #define cBuffer aMenu[16]

#define nRow aMenu[17] ; #define nCol aMenu[18]

#define lMouseHit aMenu[19] ; #define nSaveRow aMenu[20]

#define nSaveCol aMenu[21] ; #define lMessage aMenu[22]

#define cGrayColor aMenu[23] ; #define nMouseStat aMenu[24]

#define nStRow aMenu[25] ; #define nStCol aMenu[26]

#define nEnRow aMenu[27] ; #define nEnCol aMenu[28]

#define lMouVisible aMenu[29] ; #define nPass aMenu[30]

#define aSubMenu aMenu[31] ; #define aSubBlock aMenu[32]

#define nElement aMenu[33] ; #define aMenuItems aMenu[34]

#define aHotKeys aMenu[35] ; #define aMenuBlocks aMenu[36]

#define aSubItems aMenu[37] ; #define aBlockItems aMenu[38]

#define cType aMenu[39] ; #

Programming Techniques for Memory Management (Part 1) by Roger Donnay

This section written by Roger Donnay is the second part of CA-Clipper 5.x Memory Management by Roger Donnay & Jud Cole.


How you write your CA-Clipper code will greatly determine how the application uses memory. Once you understand and become grounded in the basics you will discover that memory management will become second nature to your programming style. You should never write code or use any third-party library without a basic understanding of the impact on memory usage, otherwise you can paint yourself into a memory-deficient corner that can not be easily undone.

Virtual Memory

CA-Clipper includes an automatic memory manager referred to as the Virtual Memory Manager or VMM. There is very little the CA-Clipper programmer needs to know about the VMM system other than it requires EMS memory or available disk space for the creation of swap files. After CA-Clipper uses up the available EMS for VMM, it will start creating swap files. In a network situation with no local hard drive or ram drive, this can slow down the application, therefore it is recommended that you set up your environment to allocate from 500k to 1 meg of EMS for average Clipper applications. You may monitor CA-Clipper's usage of EMS during your application by reporting the remaining EMS with the MEMORY(4) function. If the number reported falls below 100, then it is recommended that you increase available EMS by several hundred K.

If no EMS is available, then Virtual memory can be allocated from XMS memory by using a Third-party product named ClipXMS. See the section titled DOS Extenders for more information on this product.

Conventional Memory

The CA-Clipper free-pool (also sometimes referred to as the fixed-heap) is always allocated from conventional memory. I have found, from experience, that applications will perform poorly or run the risk of running out of memory completely if the amount of free-pool deteriorates to less than 50k. You can count on the fact that your application will have less free-pool available after running for awhile than it has at start-up time. The CA-Clipper MEMORY(0) function can be used to monitor the amount of memory remaining in the fixed heap. Opening databases has the most serious long-term affect on the use of conventional memory. See the section titled "The Affects of Databases on Memory Usage" for more details. To give the CA-Clipper application more conventional memory, you must start at DOS with ample conventional memory. After you load device drivers, network drivers, mouse drivers, TSR's, etc. your environment just may not have sufficient conventional memory remaining to run your Clipper application. The new generation of expanded memory managers are designed to not only create EMS/XMS memory but they also will load drivers and TSR's into the upper memory blocks (UMB) area of memory above 640k so they will not use up valuable conventional memory. I have used QEMM and 386MAX in the past to accomplish this task but have recently found that, on some systems, the MEMMAKER.EXE utility and the EMM386 driver supplied with DOS 6.x are adequate in creating a better memory environment.

SYMBOL Management

For a dynamic-overlay manager to work effectively, it is important that it load overlays at the smallest possible code-segment level, i.e., the procedure/function rather than the entire object. This requires managing the symbol table separately from the code and data, therefore, these linkers place the symbol table into the root memory area and eachfunction into a separate overlay segment.

Symbols cannot be overlayed or swapped to VMM, therefore it is important that you program in a manner consistent with producing the smallest number of symbols in your CA-Clipper-compiled objects. Here are some tips for reducing the symbol table size in your applications.

Use Constants instead of Memvars

All PRIVATE, PUBLIC and STATIC CA-Clipper memory variables are treated as "symbols". Refrain from using a memory variable if a constant is sufficient. For example, an unnecessary symbol can be eliminated by changing the code:

nEscapeKey := 27

DO WHILE INKEY() # nEscapeKey

* CA-Clipper code

ENDDO

to:

DO WHILE INKEY() # 27

* CA-Clipper code

ENDDO

or:

#define K_ESC 27

DO WHILE INKEY() # K_ESC

* Ca-Clipper code

ENDDO

Use Arrays Instead of Memvars

Every different CA-Clipper PRIVATE, PUBLIC, or STATIC memvar name creates a "symbol", whereas an array name creates only ONE symbol. The following example shows how to save considerable memory in a CA-Clipper application by reducing the symbol count with an array.

This code produces 5 symbols:

PRIVATE cName := customer->name

PRIVATE cAddress := customer->address

PRIVATE cCity := customer->city

PRIVATE cState := customer->state

PRIVATE cZip = customer->zip

@ 1,1 SAY 'Name ' GET cName

@ 2,1 SAY 'Address' GET cAddress

@ 3,1 SAY 'City ' GET cCity

@ 4,1 SAY 'State ' GET cState

@ 5,1 SAY 'Zip ' GET cZip

READ

This code produces 1 symbol:

PRIVATE aGets[5]

aGets[1] := customer->name

aGets[2] := customer->address

aGets[3] := customer->city

aGets[4] := customer->state

aGets[5] := customer->zip

@ 1,1 SAY 'Name ' GET aGets[1]

@ 2,1 SAY 'Address' GET aGets[2]

@ 3,1 SAY 'City ' GET aGets[3]

@ 4,1 SAY 'State ' GET aGets[4]

@ 5,1 SAY 'Zip ' GET aGets[5]

READ

Some programmers choose memvars over arrays because it makes their source code more readable. Source code that refers to hundreds of array elements can look very cryptic and can be hard to maintain. To reconcile this problem, use the CA-Clipper pre-processor to create a "symbolic" reference to each array element like in the following example.

STATIC aGets[5]

#define cNAME aGets[1]

#define cADDRESS aGets[2]

#define cCITY aGets[3]

#define cSTATE aGets[4]

#define cZIP aGets[5]

cNAME := customer->name

cADDRESS := customer->address

cCITY := customer->city

cSTATE := customer->state

cZIP := customer->zip

@ 1,1 SAY 'Name ' GET cNAME

@ 2,1 SAY 'Address' GET cADDRESS

@ 3,1 SAY 'City ' GET cCITY

@ 4,1 SAY 'State ' GET cSTATE

@ 5,1 SAY 'Zip ' GET cZIP

READ

Use the Same Name Memvars whenever possible

Again, every "different" PUBLIC, PRIVATE or STATIC CA-Clipper memvar in a module creates a symbol. If an object contains several procedures, use the same name for memvars even though they may not perform the same or similar functions. For example, procedure A and procedure B both need 5 memvars. If procedure A declares its memvars with 5 unique names and procedure B declares its memvars with 5 unique names, then 10 symbols are used in the linked application. To eliminate 5 symbols, make sure that procedure B assigns the same name to the memvars as procedure A. This is not possible of course, if the memvars need to be PUBLIC to both procedures and perform different functions, only if they are PRIVATE.

Use Complex Expressions instead of Memvars

The following three lines of code represents the method that most CA-Clipper programmers choose to accomplish most programming tasks. It makes sense to code this way for readability and debugging, but if you are writing a very large application, the complex expression technique can save some memory. The following three lines of code will read the disk file READ.ME into a memvar named cReadFile, save the changed code into a file named cEditFile, then write the changed code back to the disk file READ.ME.

cReadFile := MEMOREAD('READ.ME')

cEditFile := MEMOEDIT( cReadFile )

MEMOWRIT('READ.ME', cEditFile )

These three lines of code can be replaced by one complex expression which uses no symbols at all.

MEMOWRIT("READ.ME",MEMOEDIT(MEMOREAD("READ.ME")))

An additional advantage to coding this way is that less free-pool or VMM memory is used because the process of temporarily storing the text in cReadFile and cEditFile is completely eliminated. If you find that creating complex expressions such as this are unreadable and hard to maintain then use the CA-Clipper pre-processor to accomplish the same task as follows:

#define cReadFile MEMOREAD('READ.ME')

#define cEditFile MEMOEDIT(cReadFile)

MEMOWRIT('READ.ME',cEditFile)

The above code is nearly as readable as the original three lines of code but will not create any variables at compile time. Try compiling this code with your CA-Clipper compiler and use the /P switch to write the pre-processed code to a .PPO file, then look at the .PPO file to see what is actually compiled.

Use LOCALS Instead of PRIVATES

Sometimes it is just not possible or practical to write code without using symbols, so if you find yourself in this situation, CA-Clipper provides the feature of "LOCALIZING" symbols to the code segment which is currently being executed rather than placing the symbol in the main symbol table. LOCAL symbols are effectively "overlayed" because they are treated as part of the code segment rather than given a place in the main symbol table.

Not only does this save valuable memory but it also improves speed performance because the public symbol table does not need to be searched each time a LOCAL symbol is referenced in your code. Of course, if the symbol you are referencing is needed for the entire application or is used in a macro, then it must be declared as PRIVATE or PUBLIC. Symbols which are not declared at all are automatically assumed to be PRIVATE, so make sure you use the LOCAL declaration for all symbols in your code which you do not want to end up in the main symbol table. In the previous code example, the 5 PRIVATE memvars consume 110 bytes of memory, the single PRIVATE array consumes 22 bytes of memory, whereas 5 LOCAL declarations would consume 0 bytes of memory.

Use STATIC functions instead of PUBLIC functions

Every PUBLIC function and procedure name will occupy 22 bytes of the symbol table. Don't make functions and/or procedures PUBLIC unless they need to be called from anywhere within your application. STATIC procedures and functions are called only from within the source code module in which they are declared thereby eliminating the need to place the function name into the public symbol table.

Use PSUEDO functions or CONSTANTS instead of PUBLIC functions

A Psuedo-function is a function that does not really exist in the compiled code but exists only in your source code. Many programmers will use a large number of functions with different names even when each function may do something very similar like a conversion or a table lookup. I have seen a lot of code that looks like this:

SetColor( blueonwhite() + "," + redongreen() + "," + black() )

FUNCTION blueonwhite

RETURN 'B/W'

FUNCTION redongreen

RETURN 'R/G'

FUNCTION black

RETURN 'N/N'

Programmers will use this technique to make their code easy to read and maintain without understanding how this can bloat the size of the application. Probably the best technique to replace the above code would be to #define constants for each as follows:

SetColor( BLUEONWHITE+ "," + REDONGREEN + "," + BLACK )

#define BLUEONWHITE "B/W"

#define REDONGREEN "R/G"

#define BLACK "N/N"

Your existing source code may have "many" references to public functions and you don't want to risk changing all the function names to constants, or maybe the functions return something a little more complex than can be handled by a simple constant. In this case, you can #translate the functions into psuedo-functions as follows:

BEFORE:

FUNCTION BLOBget ( nPointer, nStart, nCount )

RETURN dbInfo( BLOB_GET, { nStart, nCount } )

FUNCTION BLOBput ( nPointer, xBlob )

RETURN dbInfo( BLOB_PUT, { nPointer, xBlob } )

FUNCTION BLOBExport( nPointer, cTargetFile, lMode )

RETURN dbInfo( BLOB_EXPORT, { nPointer, cTargetFile, lMode } )

AFTER:

#xTranslate BLOBget( , , ) => ;

dbInfo( BLOBGET, { , , } )

#xTranslate BLOBput( , ) => ;

dbInfo( BLOBPUT, { , } )

#xTranslate BLOBExport( , , ) => ;

dbInfo( BLOBEXPORT, { , , } )

In the BEFORE example above, three PUBLIC functions were created simply to provide a better way to call the same dbInfo() function, whereas in the AFTER example, no PUBLIC functions were needed to accomplish the exact same task. Not only did we save 3 symbol table entries but we also eliminated 3 LOCAL variables that get pushed onto the EVAL stack. Of course it must be remembered that since psuedo-functions don't actually exist they cannot be used in macros or index keys.

Don't use STATICs or PRIVATEs to pass parameters

Many CA-Clipper programmers will assign a variable as PRIVATE or STATIC so it can be accessed and changed within multiple procedures in the same source code module. Variables should be STATIC only if their value needs to be maintained throughout the program, not for the convenience of eliminating the need to pass parameters. By using the pass-by-reference symbol "@" you can change the value of LOCAL variables anywhere within your calling program as shown by the following example.

LOCAL cName, cAddress

MyFunction( @cName, @cAddress )

Return( { cName, cAddress } )

STATIC FUNCTION MyFunction ( cName, cAddress )

cName := "CA-Clipper"

cAddress := "New York, NY"

Return( nil)

The "Myth" of PUBLICs verses STATICs

There is no memory advantage to using a STATIC variable in lieu of a PUBLIC variable. Both types of variables use up permanent space in conventional memory. A STATIC variable will use space in DGROUP, whereas a PUBLIC variable will use space in the Symbol Table. In fact, there are situations in which a good case can be made for using a PUBLIC array instead of a STATIC array. Take the example of a system-wide color system in which the colors are defined in an array. If the array is STATIC, then extracting the information from the array requires a call to a PUBLIC function which exists in the same source code module as the array definition. This would actually add more symbols to the symbol table and take more processor time than if the color array were defined as PUBLIC thereby allowing direct access to the array without the need for the public function. A STATIC array may be a wiser choice in cases where the application is a library which is incorporated into another CA-Clipper application, thereby eliminating any possibility of symbol conflicts with the application.

DGScan-The DGROUP Usage Scanner For Clipper by Ian Day & Dave Pearson

By Ian Day and Dave Pearson

Background On DGROUP

(Portions of this section are taken, with the kind permission of Dark Black Software, from the MrDebug Norton Guide)

DGROUP is a 64K chunk of memory that, for a Clipper program, can be broken down into five distinct sections, each serving a specific purpose. DGROUP can be easily described as a 64K block of memory that is used by Clipper as a common area of memory that Cliper and third party products can rely on to find various items of information. There are five main parts to DGROUP:

1. The MemVar table (Dynamic)

This is used to store Clipper STATIC variables and ITEMs (declared through the ITEM API) with their contents if the contents and variable/ITEM definition can be held within 14 bytes, otherwise the variable/ITEM definition is held along with a pointer to a Virtual Memory segment where the variable/ITEM contents are held.

2. DS Available (Dynamic)

This is the amount of free memory available within the DGROUP. This amount can be seen from //INFO as DS AVAIL or from the Memory/Info window. This will be reduced during the execution of the program as the Eval Stack and the Memvar table may both 'grow' into this area.

3. Eval Stack (Dynamic)

This is where Clipper LOCAL variables are held (if the variable definition and contents take up 14 bytes or less), otherwise a pointer is stored to a Virtual memory segment that contains the variable contents.

4. CPU Stack (Fixed)

The CPU stack is used to store the function return addresses each time a new function is called, as well as local 'C' variables.

The size of the CPU stack is set when you link your program. It depends upon the linker that you are using and linker specific commands that you have in the link script that might increase or reduce the stack

5. Fixed C and ASM Data (Fixed)

This area is used by other third party libraries and other languages that do not use their own data segments and use the default data segment (otherwise known to us Clipperites as DGROUP) to store fixed pieces of text and global variables used by other languages

For example, inside Clipper, the un-recoverable error messages are stored within this section of the default data segment (DGROUP).

As you can see from the above, the DGROUP area is pretty important to the smooth running of your Clipper application, and a lack of DGROUP can cause your software to fall over with a number of internal errors.

In an effort to make sure this does not happen you should try to reduce the impact your code has on DGROUP. Many 3rd party libraries use large ammounts of DGROUP (in my experience they can bump up the fixed C and ASM data usage by quite a bit. If you need to use a number of 3rd party libraries this can become quite a problem.

However, given that your chances of changing the usage of 3rd party libraries are pretty slim, the only chance you have of making a difference is by improving your own code to reduce it's DGROUP usage.

What Can Be Done?

If you write nothing but pure Clipper code, you still have the ability to reduce your DGROUP usage. The main area you can address is your use of STATIC variables. STATIC variables are, without a doubt, a good thing, but, they come with a price.

Unlike the other variable types, each STATIC variable requires a fixed (14 byte) entry in the the MemVar table. On the surface this may not seem like a big deal. A STATIC only needs 14 bytes in DGROUP, what does it matter? To see how this matters, lets take the GETSYS.PRG code you can find in the SOURCE\SYS directory of your copy of Clipper as an example.

Close to the top of the source you will find the following code:


//
// State variables for active READ
//
STATIC sbFormat
STATIC slUpdated := .F.
STATIC slKillRead
STATIC slBumpTop
STATIC slBumpBot
STATIC snLastExitState
STATIC snLastPos
STATIC soActiveGet
STATIC scReadProcName
STATIC snReadProcLine


As you can see, there are 10 STATIC variables, each using 14 bytes in the MemVar Table. With a simple little trick we can reduce that 140 bytes of usage down to 14 bytes. Watch:


// Hold all the statics in one static.

Static _aSysStuff := { NIL, .F., NIL, NIL, NIL, NIL,; NIL, NIL, NIL, NIL }

// Translated static variable names.

#xtranslate sbFormat => _aSysStuff\[ 01 \] #xtranslate slUpdated => _aSysStuff\[ 02 \] #xtranslate slKillRead => _aSysStuff\[ 03 \] #xtranslate slBumpTop => _aSysStuff\[ 04 \] #xtranslate slBumpBot => _aSysStuff\[ 05 \] #xtranslate snLastExitState => _aSysStuff\[ 06 \] #xtranslate snLastPos => _aSysStuff\[ 07 \] #xtranslate soActiveGet => _aSysStuff\[ 08 \] #xtranslate scReadProcName => _aSysStuff\[ 09 \] #xtranslate snReadProcLine => _aSysStuff\[ 10 \]

With that simple little trick you will have reduced the DGROUP usage from 140 bytes to just 14 bytes.

If you have lots of code of your own that has a large number of file wide STATICs in a single file you can make quite a difference.

If you write your own C functions, then it's a wise move to take a look at the compiler options at your disposal, as some of these can drastically reduce the impact on DGROUP.

For example, take the following piece of code:


#include

CLIPPER YesNo( void ) // Return yes or no string { if ( _parl(1) ) // .T. passed to function? { _retc( "Yes" ); // Yes, so return "Yes" } else { _retc( "No" ); // No, so return "No" } }

Simple enough, but already we've just used 7 (or 8) bytes of DGROUP without thinking! Why? you ask. Answer: It's the strings. By default, C compilers will put constants (which means strings, doubles, floats and some arrays) into DGROUP. And when I say '7 (or 8)', it's because some compilers will align data on a WORD boundary as default, and you can usually expect the DGROUP usage to be a tad more than you thought.

You can change this by looking for a compiler option to reverse this default. For example, with the Microsoft C compilers, there is the /Gt switch and for Borland C compilers there is the -Ff switch. Given a number of 1, it will automatically put any constants larger than 1 length into a FAR segment, and out of DGROUP! This means, that no matter how much data is used by your C functions for strings and things, it won't make a difference to DGROUP.

With assembler code, we use the same principles. Don't use _DATA, _BSS, CONST or DGROUP for data storage if you can possibly help it.

So, if you have a module like this:


_DATA SEGMENT WORD PUBLIC 'DATA'
cYes DB 'Yes', 0
cNo DB 'No', 0
_DATA ENDS

DGROUP GROUP _DATA

YESNO_TEXT SEGMENT WORD PUBLIC 'CODE'

ASSUME cs:YESNO_TEXT, ds:DGROUP

YESNO PROC FAR PUSH bp MOV bp, sp

MOV ax, 1 PUSH ax CALL __parl ADD sp, 2

OR ax, ax JZ ZeroSoItsFalse

MOV ax, OFFSET cYes JMP SHORT ReturnString

ZeroSoItsFalse: MOV ax, OFFSET cNo

ReturnString: PUSH ax PUSH ds CALL __retc ADD sp, 4

POP bp RETF YESNO ENDP

YESNO_TEXT ENDS END

You've again used 7 (or 8) bytes of DGROUP without thinking, so in this case you would have to alter the segments and anything that used the data in them so that DGROUP can once again be saved:


YESNO_DATA SEGMENT WORD PUBLIC 'FAR_DATA'
cYes DB 'Yes', 0
cNo DB 'No', 0
YESNO_DATA ENDS

YESNO_TEXT SEGMENT WORD PUBLIC 'CODE' ASSUME cs:YESNO_TEXT, ds:NOTHING

YESNO PROC FAR PUSH bp MOV bp, sp

MOV ax, 1 PUSH ax CALL __parl ADD sp, 2

OR ax, ax JZ ZeroSoItsFalse

MOV ax, OFFSET cYes JMP SHORT ReturnString

ZeroSoItsFalse: MOV ax, OFFSET cNo

ReturnString: PUSH ax PUSH YESNO_DATA CALL __retc ADD sp, 4

POP bp RETF YESNO ENDP

YESNO_TEXT ENDS END

As you can see, if your application is tight on DGROUP, with the source to hand and a little bit of work you can free up some of that usage. Lets just hope you have the source code for the worst offending code in your application.

Using DGSCAN To Find DGROUP Usage

Back when I first became aware of this "problem" I started to work on my own library code and managed to reduce some of the DGROUP usage. However, working through all that code was getting to be pretty boring. That's when I decided to write DGSCAN. What I wanted was a tool that could scan OBJ and LIB files and find anything that might be taking up precious DGROUP space.

If you want to get and try out DGSCAN then pop over to Hagbard's World and grab a copy. Got it? Good.

Ok, lets take a simple example of using DGSCAN. Look at the following code:

Static xVar1
Static xVar2
Static xVar3
Static xVar4

Function Main()

? "I'm eating up DGROUP"

Return( NIL )

Compile it into an OBJ and then run DGSCAN over it:

F:\DGTEST>dgscan foo.obj

DgScan v3.00 - DGROUP Usage Scanner By Dave Pearson

File-------- Module--------------------------- Segment----------- Bytes----- FOO.OBJ FOO STATICS$ 56 FOO.OBJ *** Total *** 56

As you can see, DGSCAN has detected 56 bytes of possible DGROUP usage. This is because we have 4 STATIC variables, each with a 14 byte impact. If you applied the "array trick" we spoke about earlier you would see the usage reduced to 14 bytes.

To generate a report for all your OBJ files just do:

F:\DGTEST>dgscan *.obj


It's as simple as that, and the same goes for LIB files.

I won't say much more about DGSCAN, everything you need to know is covered by the file DGSCAN.TXT found in the DGSCAN archive. However, if you do have problems with the utility please feel free to mail me and I'll try to help you out.

The CA-Clipper Memory System - by Jud Cole

This section written by Jud Cole is the first part of CA-Clipper 5.x Memory Management by Roger Donnay & Jud Cole. If you like to have the whole book in .WRI format.

Jud Cole is the President of Blink, Inc and has been programming on microcomputers since 1979 in many languages including Assembler, C, Pascal, PL/I, dBase and CA-Clipper. After working as a consultant programmer in the early 1980's he worked at IBM for three years in their PC division doing training and support on their entire range of PC products. Following this he wrote contract database applications interfacing CA-Clipper databases to external devices such as magnetic card readers, vehicle tachographs and mechanised storage systems. During 1989 and 1990 he developed BLINKER, the first dynamic overlay linker. Since then he has been enhancing and promoting BLINKER and speaking at user groups and conferences.


The aim of this section is to impart a greater understanding of how a CA-Clipper application works internally, with a view to writing more compact and efficient CA-Clipper applications.

This section describes how the CA-Clipper 5.x Virtual Memory Manager uses conventional memory, expanded memory and disk space to store both data and CA-Clipper code. We will examine the PUBLIC, PRIVATE, LOCAL and STATIC variable classes, storage of memory variable values of different types, and how the dynamic paging system manages CA-Clipper 5.x code at application run time.

The level of expertise of the reader is expected to be medium to high, assuming an in-depth knowledge of CA-Clipper programming, and a good knowledge of PCs, networks and programming techniques in general.

Terms and definitions Conventional memory is the memory which exists on all PCs and compatibles, and typically consists of 512kb or 640kb of memory on the mother board. Due to the architecture of the early IBM PCs, the maximum amount of conventional memory is usually limited to 640kb, although certain memory managers can provide another one or two hundred kb on some machines. In theory, the maximum conventional memory on a 8088 / 8086 processor is determined by the 1 Mb address space.

Expanded memory is memory which is also accessible on all PC compatibles. Programs which are to use expanded memory have to be explicitly written to do so. Expanded memory may be provided in the form of hardware or software emulation, and in later versions of the specification, known as EMS, is limited to 32 MB. It is managed in pages, typically 16 kb in size, which may be brought into an area of conventional memory to be accessed by a program.

The currently executing program will request a particular page of expanded memory from the expanded memory manager, and will provide an address in conventional memory at which to place the page. On return from the manager, the data contained in the requested page can be read or written to as if it were permanently resident in conventional memory.

Extended memory is memory accessible by the 80286 and later processors, and exists outside of the 1 Mb range of conventional memory. These processors can access extended memory directly when running in one of their enhanced modes. When running in an 8086 emulation mode, however, a programming interface to extended memory is necessary, and a number of these have been specified. The most widely used of these interfaces is known as the XMS specification.

Software memory managers will often manage extended memory and provide both EMS and XMS programming interfaces to it for maximum versatility.

Virtual memory is a technique which has been used for many years to enable programmers to write programs requiring more memory than is directly available on the destination machine. The technique provides a simple interface to memory for storing and retrieving code and data, whilst hiding the fact that the information may be stored on one or more alternative devices until it is needed again.

The virtual memory manager, which may be implemented in hardware, software, or a combination of the two, monitors the frequency and duration of usage of the information, and decides where to keep each piece of information for maximum overall performance of the system. Typically, the least recently used information will be saved out to slower devices, while the more recently or more often used information will be kept in fast, real memory.

CA-Clipper's virtual memory manager

CA-Clipper 5.x contains its own virtual memory manager, known as the VMM, to manage the data and CA-Clipper code belonging to the application. By default, the CA-Clipper VMM allocates all available conventional memory and up to 8 MB of expanded memory for this purpose. In addition, if all the available memory has been used up, the VMM will swap out data, but not code, to a temporary file on disk.

CA-Clipper currently makes no direct use of extended memory, so if the application will be running on a 386 PC or above, then obtaining a memory manager such as QEMM, 386MAX or the one supplied with MS DOS 6.0 will be a good investment.

Once a CA-Clipper application has loaded into memory and started executing, the application allocates the remaining real memory according to parameters set with the CLIPPER environment variable or the // command line options. The format of these is the same, and consists of the //, the letter or group of letters denoting the area, e.g. E for EMS, a ':' and a number indicating the size in Kb to be used for that area.

The parameters controlling allocation of memory are X:nnn and E:nnn. The X parameter specifies how much conventional memory to eXclude from use by CA-Clipper, and takes a value from 0 to 256 kb. The E parameter specifies how much expanded memory to allocate to the VMM, and takes a value from 0 to 8192 kb.

For example : TEST //E:1000

which would limit CA-Clipper to using 1 Mb of expanded memory.

It is worth noting in passing that the default of all available EMS up to a maximum of 8MB, or the E value if one is specified, is allocated to the VMM in one block at the start of the application. This means it is not available to any other part of the system until the application terminates and the memory is freed. In addition, the application could run out of conventional memory if there is too much EMS available to it, since a table of proportional size to the amount of EMS used is allocated in conventional memory. Depending on the amount of data manipulated by the application, a suitable maximum value may be 1000 to 2000, representing 1 - 2 Mb of expanded memory.

The other CA-Clipper parameters relevant to the VMM are the SWAPPATH:'path' and SWAPK:nnn parameters. If the application's conventional memory and EMS memory is fully utilised then the VMM will create a temporary swap file in the directory indicated by the SWAPPATH parameter, or in the current directory if no SWAPPATH is specified. This disk file will be used to store the least recently used data owned by the VMM, and will gradually increase in size until either the application has terminated and the file deleted, or the size limit set by the SWAPK parameter has been reached. The default size limit for the swap file if no SWAPK parameter is specified is 8 MB.

Virtual memory as managed by the VMM is allocated in segments, each of which may contain from 1K to 64K of data. When memory is allocated from the VMM, instead of returning a pointer to real memory it returns a form of segment number to identify the segment, in the same way as DOS returns a handle when a file is opened. Whenever the data within the segment is needed, a request is made to the VMM to return the current location of the segment in real memory where it can be read or written to.

Initially all the segments will be located in real memory, and because each segment is movable, real memory can be organised efficiently by filling up the gaps as segments are freed. Once real memory fills up, the VMM will swap out least recently used segments to EMS if it is available, or disk if not, to make room for new segments. If those segments are used at a later stage in the program, the VMM will swap out other segments to make room and bring the original segments back in, in the same way as an overlay manager manipulates code overlays.

CA-Clipper 5.x also contains a special type of memory manager designed to manage complex data values such as character strings and arrays. The CA-Clipper 5.x object memory is called the Segmented Virtual Object Store (SVOS). SVOS uses virtual memory managed by the VMM to store data values, including character strings, arrays, and dynamically created (macro-compiled) code blocks.

SVOS provides two important functions beyond the basic capabilities offered by the VMM, memory compaction and garbage collection.

Memory compaction consists of automatically compacting stored values on an ongoing basis. This eliminates fragmentation of the virtual memory and reduces swapping, since each segment can be fully utilised before requesting further segments.

Some CA-Clipper 5.x values (e.g., arrays) may be referred to by several program variables or array elements at the same time. The garbage collection routine automatically reclaims space occupied by values which are no longer accessible through any variable or array. By default, this occurs in background when CA-Clipper is in an idle state, e.g. waiting for keyboard input.

The real memory remaining to the VMM is set up as a swap area to bring swapped out pages of data into memory for use in the CA-Clipper program. When a RUN command is performed, as much of the top of the swap space as possible is cleared and returned back to DOS to be combined with the X area, and then the RUN command is issued. In this way more memory is freed up for RUN commands than would have been with Summer '87, although the exact amount will depend on the size and usage of the lower end of the swap space.

CA-Clipper's symbol table

The CA-Clipper language, and the dBase language on which CA-Clipper was originally based, is a dynamic language with a number of very powerful constructs. These allow and cause certain functions normally performed by the compiler to be postponed until run time, such as setting the type of variables and using macros to create new variables not known at compile time.

Because of this dynamic nature, at run time CA-Clipper requires more information about variables and procedures than traditonal lower level languages such as Pascal, C and Modula 2. Some of this information is available at compile and link time, such as the name of the variable, but some of it, such as its type, will only be available once the application has started executing.

For these reasons, each CA-Clipper .OBJ file is created with a symbol table of 16 bytes per symbol, and all code in the .OBJ file refers to that symbol table. At run time the symbol table entry is used to point to the control information and value or code for the symbol. The symbol table is created in its entirety in the root of the application, and can grow to upwards of 64 kb, so it can significantly affect the amount of conventional memory required by the application. This is why even 100% overlayed applications grow when code is added.

CA-Clipper 5.x introduced static and local variables to the language to encourage better and more efficient coding practices. Another important benefit is that these classes of variables do not require a symbol table entry as they cannot be accessed via macros. Changing as many PUBLIC and PRIVATE variables as possible to LOCAL or STATIC variables can therefore significantly reduce the amount of conventional memory required.

The major linkers now available remove the duplicate symbols from the symbol tables in the various .OBJ and .LIB files at link time, creating one large consolidated symbol table. This process, known as symbol table compression, can significantly reduce the run time memory requirement of the .EXE, leaving more memory for the application's data and overlays. All the duplicates are removed except the symbols belonging to procedures declared as static, since these are local to each .OBJ and will have different code associated with each occurrence of the symbol.

It is worth noting that prior to link time symbol table compression, the only way to reduce the number of duplicate symbols was to minimise the number of .OBJ files, but this is no longer necessary.

CA-Clipper code

When compiled, each CA-Clipper procedure or function in the .OBJ file has a separate unit know as a segment, which consists of a small Assembly language header and a string of tokens. The header simply consists of pointers to the CA-Clipper symbol table and the tokenised code and a call to the CLIPPER.LIB procedure __PLANKTON. The tokenised code represents calls to functions within the CA-Clipper library and parameters to those functions. At application run time when the procedure or function is called the __PLANKTON procedure processes these tokens sequentially and performs the appropriate library calls with the parameters held in the tokens. Each token is usually only one byte long, with parameters varying in length, e.g. a real number will take up 8 bytes and a character string will be stored as the length followed by the string. Tokens may also refer to symbols in the symbol table described above, rather than referring to absolute locations, so each reference to a variable will consist of a two byte symbol number.

For example, in the code :

FUNCTION T A = B + C

we would have a symbol table containing :

T A B C

and the tokenised code would consist of (in simplified terms) :

Take symbol 2 (B) Take symbol 3 (C) Add them together Store result in symbol 1 (A)

This tokenised approach has a number of advantages over true compiled code, with only a neglible cost in performance. The code produced is very compact, for example taking only three bytes for a procedure call, as opposed to five for a direct call. It is also very self contained. All external references go via the symbol table, so operations such as incremental linking are made significantly easier. This approach also makes it possible to use the dynamic paging system described below for faster overlayed applications with lower memory requirements.

The size overhead of a simple CA-Clipper compiled .EXE is actually made up of the runtime routines from the CLIPPER.LIB which are called by the processing of the tokens. The apparently large size of even a "Hello world" type program is due to the potential for macro operations, which could execute just about any CA-Clipper command from even a two line program.

Instances of variables and their values

In conventional languages the scope, size and type of a named variable is known at compile time, so the exact amount of space can be reserved for it in memory at run time. This memory will always be used to store the value, no matter how often the value is changed.

The remaining memory above the program's .EXE image is usually known as the heap and is managed by a heap manager, which will allocate blocks of memory of varying size to the program as and when requested. Space for data allocated dynamically at run time, for constructs such as linked lists or buffers, whose sizes are not known at compile or link time, will be allocated from and returned to this heap.

With CA-Clipper, determination of the type and size of all variables and the scope of public and private variables is left until run time, so a more complicated mechanism for storing the values of variables is required.

CA-Clipper 5.x offers several different storage classes for program variables, depending on how they are declared and used in the program. LOCAL and STATIC variables are stored in a dedicated area of real memory, as described below. PRIVATE and PUBLIC variables, known as MEMVAR variables, are created and destroyed dynamically while a program is running, and are stored in VM segments.

For performance reasons, these segments remain locked in real memory during most operations except memory intensive operations and RUN commands. Each MEMVAR uses 20 bytes in a VM segment, so converting PRIVATE and PUBLIC variables to LOCAL and STATIC variables can reduce memory requirements for some applications.

At run time, each instance of a variable is allocated a value, which is represented internally as a data structure called a VALUE. The contents and format of a VALUE differ depending on the type of data it represents. Simple data, such as integers, are stored directly into the VALUE. Larger items, or data of variable length such as strings or arrays, have a "reference" to the string or array stored in the VALUE, and the actual data is stored elsewhere. Internally, CA-Clipper is organized as a stack based machine which uses an area of memory called the Eval Stack to contain temporary variables such as function parameters, intermediate results of expressions and local variables. The Eval Stack is simply a contiguous group of VALUEs that are accessed as a stack, in the same way as the processor stack is used by C programs.

For example, in a CA-Clipper function call, parameters are pushed onto the Eval Stack before the function is executed. The function operates on the top-most items in the Eval Stack and produces a result. After the function completes, the parameter values are popped from the Eval Stack and replaced with the function result.

Each entry in the Eval Stack, i.e. each VALUE, occupies 14 bytes, and for complex data types such as character strings, arrays and code blocks there will be an additional memory requirement handled by the VMM where the actual value is stored.

The Eval Stack is allocated from the default data segment, defined as the start of the group DGROUP, when the program starts executing, so initialisation will fail if DGROUP is too full. This is not usually a problem with pure CA-Clipper applications, but if a number of third party libraries are linked in to the application it may possibly fill up unless they have avoided storing data in DGROUP. The number of kb remaining in DGROUP for CA-Clipper's use can be examined by executing the program, with the //INFO parameter, and the amount of conventional and expanded memory available will be displayed at the same time.

LOCAL variables are the simplest variables, and are allocated as locations within the Eval Stack to store their VALUEs. To manipulate a LOCAL variable, the system simply copies the variable's VALUE from one position in the Eval Stack to another.

Local variables are visible only within the current procedure or function, and are created automatically each time the procedure in which they were declared begins executing. When that procedure terminates through a return, all it's LOCALs are removed from the Eval Stack and any associated VMM memory freed up.

STATIC variables are similar to LOCAL variables, but have a duration of the lifetime of the application. Because of their permanence, they are allocated as fixed locations at one end of the Eval Stack, but are manipulated in the same way as LOCAL variables simply by copying their VALUEs.

This means that every STATIC variable in the system also requires 14 bytes on the Eval Stack in DGROUP, which is another reason for C and ASM programmers to avoid storing data in DGROUP.

PRIVATE and PUBLIC variables are more complex than LOCAL or STATIC variables because in addition to an associated VALUE they also have a name which may be referred to during execution of the program via a macro or its equivalent. MEMVAR variables are allocated locations for their VALUEs in dedicated VM segments and these locations are stored with their names in the symbol table.

When a MEMVAR is manipulated, the symbol table entry is used to point to the VALUE which can then be placed on the Eval Stack in the normal way. FIELD variables differ from the other storage classes because they have no memory location at all, since their values are stored in a database record buffer. To manipulate a FIELD, the system generates a request to the file's database driver, which then creates an appropriate VALUE to be manipulated on the Eval Stack.

Arrays

An array VALUE contains a reference to the array rather than an actual value, so when an array is assigned to a variable, the system simply overwrites the variable's VALUE with a new VALUE containing a reference to the array. The array itself is simply a group of VALUEs stored in virtual memory, where each element of the array is a VALUE. Any VALUE can contain another reference, so multidimensional arrays are created by having each element refer to another array rather than have an absolute value. When values are assigned to array elements, the VALUE for that element is updated. When an array is assigned to another variable, only a copy of the VALUE referring to the array is made, and the array data itself is not duplicated.

Character Values

A character string VALUE contains a reference to the character data, which is stored elsewhere in the VM. As with arrays, assigning a character value to a variable simply overwrites the variable's VALUE with a new VALUE containing a reference to the character data.

In a similar way to arrays, assigning a character value from one variable to another simply duplicates the VALUE (i.e., the reference to the data). The character data itself is not duplicated.

This reference-based memory management technique is the same for strings, arrays, and code blocks. CA-Clipper's garbage collector monitors references to objects, and when there are no longer any references to a particular piece of data, the space occupied by that data is automatically reclaimed.

Macros

During program execution, when a macro is evaluated to the name of a variable or procedure, the symbol table is searched to find the requested name. Once the name is found, CA-Clipper follows the pointer in the symbol table to the VALUE where all the general information about the symbol is actually stored. The VALUE will indicate whether the procedure or variable being referenced has been defined, and CA-Clipper checks this before continuing any further. If it is undefined and is not a variable being created, CA-Clipper immediately returns an appropriate error - "undefined function" for procedures or functions, and "variable does not exist" for variables. If the procedure or function has been defined correctly, then the VALUE will contain a pointer to the program code to execute for that procedure, and control can be transferred to the procedure.

The remaining case of creating a new variable is handled by adding a new entry to the end of the symbol table. This new entry will have the name of the variable filled in, along with a pointer to a VALUE for the symbol, and will be used from then on to refer to the variable.

Both Summer '87 and CA-Clipper 5.x provide other mechanisms to avoid the creation of these dynamically named variables in the majority of circumstances, such as using an array of elements to store the values, or using code blocks in 5.x. These alternative mechanisms should be used wherever possible, if only because macro operations are inherently very slow, as each name in the symbol table has to be checked until a match is found before execution can continue.

If the use of a macro cannot be avoided, but the name to be created will be one of a known set, then these names should be mentioned explicitly somewhere in one of the programs. The code does not ever have to be executed, but just using the names causes them to be added to the symbol table at compile time, thus avoiding the above situation.

Code blocks

Code blocks are represented internally as strings of tokenised code, in the same way as normal procedures and functions. When a code block is assigned to a variable at run time, a pointer to the tokens making up the code block is stored in the variable, along with information pertaining to the currently active procedure.

Because the code block consists of normal tokens, it will include references to the symbol table, so the equivalent symbol table must be available when the code block is actually evaluated. This is one of the reasons why it will prove difficult (but not impossible) to save code blocks in a database from one application and restore and evaluate them at a later time in the same or another application.

CA-Clipper's dynamic paging system When linked with BLINKER or .RTLink, CA-Clipper 5.x performs its own form dynamic overlaying of compiled CA-Clipper code, which results in extremely fast, memory efficient execution of the code.

During linking all CA-Clipper modules are broken down into pages of 1 kb in size. These pages are stored either in the executable file or in separate overlay files. The manipulation of overlays in these 1 kb pages removes the effect the size of compiled functions or modules has on the memory required to load the overlay. Large modules are broken into multiple pages, and small functions are grouped together in a single page.

At execution time, CA-Clipper 5.x's dynamic overlay manager loads pages based on information embedded in the .EXE by the linker. The dynamic pages are loaded into VM (Virtual Memory) segments, allowing the VMM to manage the overlay pages on a competitive basis with other uses of memory such as the application data.

The paging architecture allows the system to discard low-use sections of code even if the code is still active, and reload it only when control returns to that piece of code. Code pages which are being heavily used are maintained in memory by the VMM's LRU swapping policy.

When possible, the VMM will place dynamic overlay pages in expanded memory, reducing overlay reads. Overlay pages are never written to the VMM disk swap file, however. If a VM segment containing an overlay page is to be removed from memory altogether, it is simply discarded. If it is needed subsequently, it is re-read from the overlay file. In addition to virtual memory, the dynamic overlay manager uses a dedicated area of real memory to cache the most active dynamic overlay pages. This page mechanism is made possible by the nature of the CA-Clipper code. As explained before, it is not actually code but a series of tokens which are processed at run time. This means that the __PLANKTON procedure from CLIPPER.LIB which is processing the tokens can detect when it has reached the end of a page and request the next one to be loaded. All CA-Clipper code is therefore overlayable, so there are no restrictions on which CA-Clipper .OBJs can be placed in the overlay area. It should be noted that linkers which use the dynamic paging mechanism of CA-Clipper 5.x automatically overlay ALL CA-Clipper code unless directed otherwise.

Welcome to Clipper... Clipper... Clipper


In 1997, then using Delphi 3, I had already created 32-bits Windows applications for HRIS, ERP and CRM. In 2007, using Ruby on Rails, an AJAX powered CRM site running on Apache & MySQL was created and I am now using Visual Studio .Net 2008 to create web-based projects and Delphi 7 for Win32 applications using SQL2005 & DBFCDX.

So, why then am I reviving the Original Clipper... Clipper... Clipper via a Blog as CA-Clipper is a programming language for the DOS world ? Believe it or not, there are still some clients using my mission-critical CA-Clipper applications for DOS installed in the late 80's and up to the mid 90's. This is testimony to CA-Clipper's robustness as a language :-)

With the widespread introduction of Windows 7 64-bits as the standard O/S for new Windows based PCs & Notebooks, CA-Clipper EXE simply will not work and it has become imperative for Clipper programmers to migrate immediately to Harbour to build 32/64 bits EXEs

Since 28th January 2009, this blog has been read by 134,389 (10/3/11 - 39,277) unique visitors (of which 45,151 (10/3/11 - 13,929) are returning visitors) from 103 countries and 1,574 cities & towns in Europe (37; 764 cities), North America (3; 373 cities) , Central America & Caribeans (6; 13 cities), South America(10; 226 cities), Africa & Middle-East (12; 44 cities) , Asia-Pacific (21; 175 cities). So, obviously Clipper is Alive & Well : -)


TIA & Enjoy ! (10th October 2012, 11:05; 13th November 2015)


Original Welcome Page for Clipper... Clipper... Clipper

This is the original Welcome Page for Clipper... Clipper... Clipper, which I am republishing for historical and sentimental reasons. The only changes that I have made was to fix all the broken links. BTW, the counter from counter.digits.com is still working :-)

Welcome to Chee Chong Hwa's Malaysian WWW web site which is dedicated to Clipperheads throughout the world.

This site started out as a teeny-weeny section of Who the heck is Chee Chong Hwa ? and has graduated into a full blown web site of more than 140 pages (actually hundreds of A4 size pages) ! This is due to its growing popularity and tremendous encouragements from visiting Clipperheads from 100 countries worldwide, from North America, Central America, Caribbean, South America, Europe, Middle-East, Africa and Asia-Pacific. Thanx Clipperheads, you all made this happen !


What is Clipper ?

You may ask, what is this Clipper stuff ? Could Clipper be something to do with sailing as it is the name of a very fast sailing American ship in the 19th century ?

Well, Clipper or to be precise, CA-Clipper is the premier PC-Software development tool for DOS. It was first developed by Nantucket Corporation initially as a compiler for dBase3+ programs. Since then, CA-Clipper has evolved away from its x-base roots with the introduction of lexical scoping & pre-defined objects like TBrowse. As at today, the most stable version ofClipper is 5.2e while the latest version, 5.3a was introduced on 21 May 1996.

As at 11th November, 1996, an unofficial 5.3a fixes file was made available by Jo French. See the About CA-Clipper 5.3a section for more details. BTW, Jo French uploaded the revised 5.3a fixes file on 20th November, 1996.

Latest News

The latest news is that CA has finally released the long-awaited 5.3b patch on 21 May, 1997.

For 5.3b users, you must a take a look at Jo French's comments on unfixed bugs in 5.3b.

BTW, have you used Click ? If you're a serious Clipperprogrammer and need an excellent code formatter, Click is a natural choice. How to get it ? Simple, access Phil Barnett's site via my Cool Clipper Sites.

32-bits Clipper for Windows ?

Have you tried Xbase ++ ? Well, I have and compared to Delphi (my current Windows programming tool of choice), I'm still sticking to Delphi.

Anyway, you should visit the Alaska Home Page. Give it a chance and then draw your own conclusions !.

The Harbour Project

Is this the future of Xbase ? Take a look at at the Harbour Project

You are Visitor # ...

According to counter.digits.com, you are visitor since 3 June 1996.

If you like or dislike what you see on this website, please drop me a line by clicking the email button at the bottom of this page or better still, by filling out the form in my guest book. If you are not sure what to write,click here to take a look at what other Clipperheads have to say.