I made this widget at MyFlashFetish.com.

Saturday, April 9, 2011

PSEUDOCODE

What is pseudocode?
Pseudocode consists of short, English phrases used to explain specific tasks within a program's algorithm. Pseudocode should not include keywords in any specific computer languages. It should be written as a list of consecutive phrases. You should not use flowcharting symbols but you can draw arrows to show looping processes. Indentation can be used to show the logic in pseudocode as well. For example, a first-year, 9th grade Visual Basic programmer should be able to read and understand the pseudocode written by a 12th grade AP Data Structures student. In fact, the VB programmer could take the other student's pseudocode and generate a VB program based on that pseudocode.

Why is pseudocode necessary?
The programming process is a complicated one. You must first understand the program specifications, of course, Then you need to organize your thoughts and create the program. This is a difficult task when the program is not trivial (i.e. easy). You must break the main tasks that must be accomplished into smaller ones in order to be able to eventually write fully developed code. Writing pseudocode WILL save you time later during the construction & testing phase of a program's development.

How do we write pseudocode?
First you may want to make a list of the main tasks that must be accomplished on a piece of scratch paper. Then, focus on each of those tasks. Generally, you should try to break each main task down into very small tasks that can each be explained with a short phrase. There may eventually be a one-to-one correlation between the lines of pseudocode and the lines of the code that you write after you have finished pseudocoding.

It is not necessary in pseudocode to mention the need to declare variables. It is wise however to show the initialization of variables. You can use variable names in pseudocode but it is not necessary to be that specific. The word "Display" is used in some of the examples. This is usually general enough but if the task of printing to a printer, for example, is algorithmically different from printing to the screen, you may make mention of this in the pseudocode. You may show functions and procedures within pseudocode but this is not always necessary either. Overall, remember that the purpose of pseudocode is to help the programmer efficiently write code. 

Therefore, you must honestly attempt to add enough detail and analysis to the pseudocode. In the professional programming world, workers who write pseudocode are often not the same people that write the actual code for a program. In fact, sometimes the person who writes the pseudocode does not know beforehand what programming language will be used to eventually write the program.

Quick Tip: Configuring Eclipse to Run for the First Time


For Java developers, Eclipse is a convenient tool to develop software in broad range of programming languages. This tip is intended to help fellow developers who never use Eclipse before but are eager to start developing applications using Eclipse IDE.
Even though Eclipse is cross-platform, there is no guarantee that the tip provided in this post will also work for version of software and environment other than ones described below:
·                                 OS: Windows
·                                 Eclipse version: Eclipse 3.5 Galileo
·                                 JDK version: Java EE 6 SDK

Before we proceed with the configuration, you should have downloaded Eclipse from its homepage. If you haven’t installed Java EE SDK, you also need to download the JDK from Sun’s Java download page (make sure you download JDK with Java EE package). Eclipse requires JDK to run. It is advised to install Java EE SDK before extracting the zip file of Eclipse.

Java EE SDK installation is straightforward. You will be provided a standard wizard and you can simply go through all the steps. In the following picture, you can see the wizard is installing Glassfish web container in the disk (the installer is in Korean, though). One thing you should pay attention to is the port number of the default web server shipped along with the JDK. If you plan to use another web container like Tomcat while working with Eclipse, it is better to avoid using port number 8080 as the default web server port. Otherwise, you will be busy starting and stopping service of each web server.

After saving the modification, click eclipse.exe or the shortcut. Voila! You are now ready to create your first software project using Eclipse.

Sample PHP Application: A Simple PHP Command Line-based File Generator


PHP is a great scripting language to build web applications. Despite the sluggish improvement and development towards a more architecturally robust, more feature-rich, and less quick-and-dirty programming language recently, I still love coding bytes in PHP. Some people may think of PHP as the language for programming the web quickly. Write some HTML, embed some javascript, add some CSS, put some PHP code, and voila… a dynamic web page is created. I won’t praise how good PHP is for developing a web application. Several companies may have done that. Name Facebook and Yahoo as examples. With some optimization to native PHP codes, both companies have shown how to use the language to cater to millions of users and run a serious business.
In this post, I’d like to highlight another feature of PHP, the command line interface (CLI). PHP CLI can be an alternative to some administrative tasks. Linux users may have been familiar with shell scripting for carrying out system management and configuration tasks. So, why must PHP? The answer is portability. The same PHP code should work not only on Linux but also on Windows. Some critics may argue that other languages may also have answer for portability. I concur to that criticism while at the same time emphasizing PHP as another viable option.
The application to be shown immediately is a file generator. It will create a file of random content with size specified by user when executing the file. The content consists of alphanumeric ASCII characters that are picked randomly. No newline are included in the file. Yet, user can tweak the application to make the generated output file also include newline.
This application can be useful for those who are learning how to write a command line-based PHP utility and also for others who want to conveniently generate files of various sizes for workload testing purpose.
Source code and snapshots are provided below:
a. Source code: filegenerator.php


If you want to play with this simple application, you can download it directly from the following link:



IMPORTANCE OF TECHNOLOGY IN BUSINESS

Technology plays a vital role in business. Over the years businesses have become dependent on technology so much so that if we were to take away that technology virtually all business operations around the globe would come to a grinding halt. Almost all businesses and industries around the world are using computers ranging from the most basic to the most complex of operations. In chapter 15, we learn about how computer technology can introduce new ways for business to compete with each other. Technology make changes such as create new products, enterprise and new customer and also supplier relationship. 





Technology played a key role in the growth of commerce and trade around the world. It is true that we have been doing business since time immemorial, long before there were computers; starting from the simple concept of barter trade when the concept of a currency was not yet introduced but trade and commerce was still slow up until the point when the computer revolution changed everything. Almost all businesses are dependent on technology on all levels from research and development, production and all the way to delivery





Small to large scale enterprises depend on computers to help them with their business needs ranging from Point of Sales systems, information management systems capable of handling all kinds of information such as employee profile, client profile, accounting and tracking, automation systems for use in large scale production of commodities, package sorting, assembly lines, all the way to marketing and communications. It doesn't end there, all these commodities also need to be transported by sea, land, and air. Just to transport your commodities by land already the use of multiple systems to allow for fast, efficient and safe transportation of commodities. 



Without this technology the idea of globalization wouldn't have become a reality. Now all enterprises have the potential to go international through the use of the internet. If your business has a website, that marketing tool will allow your business to reach clients across thousands of miles with just a click of a button. This would not be possible without the internet. Technology allowed businesses to grow and expand in ways never thought possible. 


The role that technology plays for the business sector cannot be taken for granted. If we were to take away that technology trade and commerce around the world will come to a standstill and the global economy would collapse. It is nearly impossible for one to conduct business without the aid of technology in one form or another. Almost every aspect of business is heavily influenced by technology. Technology has become very important that it has become a huge industry itself from computer hardware manufacturing, to software design and development, and robotics. Technology has become a billion dollar industry for a number of individuals.



The next time you browse a website to purchase or swipe a credit card to pay for something you just bought, try to imagine how that particular purchase would have happened if it were to take place without the aid of modern technology. That could prove to be a bit difficult to imagine. Without all the technology that we are enjoying now it would be like living in the 60's again. No computers, no cellular phones, no internet. That is how important technology is in business.



Sunday, March 27, 2011

DATA INTEGRITY ISSUES


Introduction

Data integrity issues are common in relational databases, especially in operational (OLTP) systems. These issues are typically fixed by ETL (Extraction, Transformation and Load) jobs that load the data into a data warehouse. However it is not uncommon to have some integrity issues even in data warehouses.
SQL Server 2005 Analysis Services supports cubes built directly from operational data stores, and it offers some sophisticated controls to manage the data integrity issues inherent in such systems. Database administrators can greatly simplify their cube management tasks by exploiting these controls.

Types of Data Integrity Issues

In this section, we will identify some of the common data integrity issues. We will use the following relational schema for our discussions:
  • The sales fact table has a foreign key product_id that points to the primary key product_id in the product dimension table.
  • The product dimension table has a foreign key product_class_id that points to the primary key product_class_id in the product_class dimension table.
    ms345138.as2k5dataintegrity_01(en-US,SQL.90).gif
    Figure 1. Relational schema

Referential Integrity

Referential integrity (RI) issues are the most common of data integrity issues in relational databases. An RI error is essentially a violation of a foreign key–primary key constraint. For example:
  • The sales fact table has a record with a product_id that does not exist in the product dimension table.
  • The product dimension table has a product_class_id that does not exist in the product_class dimension table.

NULL Values

Although NULL values are common and even valid in relational databases, they need special treatment in Analysis Services. For example:
  • The sales fact table has a record with NULL values in store_salesstore_cost and unit_sales. These could be interpreted as a transaction with zero sales, or as if the transaction did not exist. MDX query results (NON EMPTY) would differ depending on the interpretation.
  • The sales fact table has a record with a NULL value in product_id. Although this is not an RI error in the relational database, it is a data integrity issue that Analysis Services needs to handle.
  • The product table has a record with a NULL value in product_name. Since this column is providing member keys or names in Analysis Services, the NULL value could be preserved, converted to an empty string, etc.

Data Integrity Controls

In this section, we discuss the various controls that Analysis Services offers to database administrators for dealing with data integrity issues. Note that these controls are not mutually independent. For example, Null Processing is dependent on Unknown Member and Error Configuration is dependent on Null Processing and Unknown Member.

Unknown Member

The Dimension object has a property called UnknownMember that takes three possible values—None, Hidden, Visible. When UnknownMember=Hidden/Visible, the Analysis Server automatically creates a special member called the Unknown Member in every attribute of the dimension. UnknownMember=Hidden indicates that the unknown member will be hidden from query results and schema rowsets. The default value of UnknownMember is None.
The UnknownMemberName property can be used to specify a meaningful name for the unknown member. The UnknownMemberTranslations property can be used to specify localized captions for the unknown member.
Figure 2 shows the Product dimension with UnknownMember=Visible and UnknownMemberName="Invalid Product".
ms345138.as2k5dataintegrity_02(en-US,SQL.90).gif
Figure 2. Product dimension

Null Processing

The DataItem object is used in the Analysis Services DDL to specify metadata about any scalar data item. This includes:
  • Key column(s) of an attribute
  • Name column of an attribute
  • Source column of a measure
The DataItem object contains many properties including the following:
  • DataType
  • DataSize
  • NullProcessing
  • Collation
The NullProcessing property specifies what action the server should take when it encounters a NULL value. It can take five possible values:
  • ZeroOrBlank—This tells the server to convert the NULL value to a zero (for numeric data items) or a blank string (for string data items). This is how Analysis Services 2000 handles NULL values.
  • Preserve—This tells the server to preserve the NULL value. The server has the ability to store NULL just like any other value.
  • Error—This tells the server that a NULL value is illegal in this data item. The server will generate a data integrity error and discard the record.
  • UnknownMember—This tells the server to interpret the NULL value as the unknown member. The server will also generate a data integrity error. This option is applicable only for attribute key columns.
  • Default—This is a conditional default. It implies ZeroOrBlank for dimensions and cubes, and UnknownMember for mining structures and models.
Note that the NullProcessing options Error and UnknownMember generate data integrity errors, but the others do not.
The following picture shows the DataItem editor for the key columns of a dimension attribute.
ms345138.as2k5dataintegrity_03(en-US,SQL.90).gif
Figure 3. DataItem Collection Editor

Error Nomenclature

Before we discuss the Error Configuration control, we need to clearly define the different types of data integrity errors that the server can encounter. We have already learned about two of them in the previous section on Null Processing. Following is the complete list:
  • NullKeyNotAllowed—This error is generated when an illegal NULL value is encountered and the record is discarded (when NullProcessing = Error).
  • NullKeyConvertedToUnknown—This error is generated when a NULL key value is interpreted as the unknown member (when NullProcessing = UnknownMember).
  • KeyDuplicate—This error is generated only during dimension processing when an attribute key is encountered more than once. Since attribute keys must be unique, the server will discard the duplicate records. In most cases, it is acceptable to have this error. But sometimes it indicates a flaw in the dimension design, leading to inconsistent relationships between attributes.
  • KeyNotFound—This is the classic referential integrity error in relational databases. It can be encountered during partition as well as dimension processing.

Error Configuration

The ErrorConfiguration object is central to the management of data integrity errors. The server comes with a default error configuration (in the msmdsrv.ini config file). The error configuration can also be specified on the database, dimension, cube, measure group and partition. In addition, the error configuration can also be overridden on the Batch andProcess commands.
The ErrorConfiguration object specifies how the server should handle the four types of data integrity errors. It has the following properties:
  • KeyErrorLogFile—This is the file to which the server will log the data integrity errors.
  • KeyErrorLimit (Default=zero)—This is the maximum number of data integrity errors that the server will allow before failing the processing. A value of -1 indicates that there is no limit.
  • KeyErrorLimitAction (Default=StopProcessing)—This is the action that the server will take when the key error limit is reached. It has two options:
    • StopProcessing—tells the server to fail the processing.
    • StopLogging—tells the server to continue processing but stop logging further errors.
  • KeyErrorAction (Default=ConvertToUnknown)—This is the action that the server should take when a KeyNotFound error occurs. It has two options:
    • ConvertToUnknown—tells the server to interpret the offending key value as the unknown member.
    • DiscardRecord—tells the server to discard the record. This is how Analysis Services 2000 handles KeyNotFound errors.
  • NullKeyNotAllowed (Default=ReportAndContinue)
  • NullKeyConvertedToUnknown (Default=IgnoreError)
  • KeyDuplicate (Default=IgnoreError)
  • KeyNotFound (Default=ReportAndContinue)—This is the action that the server should take when a data integrity error of this type occurs. It has three options:
    • IgnoreError tells the server to continue processing without logging the error or counting it towards the key error limit.
    • ReportAndContinue tells the server to continue processing after logging the error and counting it towards the key error limit.
    • ReportAndStop tells the server to log the error and fail the processing immediately (regardless of the key error limit).
Note that the server always executes the NullProcessing rules before the ErrorConfiguration rules for each record. This is important since NULL processing can produce data integrity errors that the ErrorConfiguration rules must then handle.
The following picture shows the ErrorConfiguration properties for a cube in the properties panel.
ms345138.as2k5dataintegrity_04(en-US,SQL.90).gif
Figure 4. Properties panel

Scenarios

In this section, we will discuss various scenarios involving data integrity issues and show how the controls described in the previous section can be used to address them. We will continue to use the relational schema specified earlier.

Referential Integrity Issues in Fact Table

The sales fact table has records with product_id that does not exist in the product dimension table. The server will produce a KeyNotFound error during partition processing. By default, KeyNotFound errors are logged and counted towards the key error limit, which is zero by default. Hence the processing will fail upon the first error.
The solution is to modify the ErrorConfiguration on the measure group or partition. Following are two alternatives:
  • Set KeyNotFound=IgnoreError.
  • Set KeyErrorLimit to a sufficiently large number.
The default handling of KeyNotFound errors is to allocate the fact record to the unknown member. Another alternative is to set KeyErrorAction=DiscardRecord, to discard the fact table record altogether.

Referential Integrity Issues in SnowFlaked Dimension Table

The product dimension table has records with product_class_id that do not exist in the product_class dimension table. This is handled in the same way as in the previous section, except that the ErrorConfiguration on the dimension needs to be modified.

NULL Foreign Keys in Fact Table

The sales fact table has records in which the product_id is NULL. By default, the NULLs are converted to zero that is looked up against the product table. If zero is a valid product_id, then the fact data is attributed to that product (probably not what you want). Otherwise a KeyNotFound error is produced. By default, KeyNotFound errors are logged and counted towards the key error limit that is zero by default. Hence the processing will fail upon the first error.
The solution is to modify the NullProcessing on the measure group attribute. Following are two alternatives:
  • Set NullProcessing=ConvertToUnknown. This tells the server to attribute the records with NULL values to the unknown member "Invalid Product". This also producesNullKeyConvertedToUnknown errors, which are ignored by default.
  • Set NullProcessing=Error. This tells the server to discard the records with NULL values. This also produces NullKeyNotAllowed errors that, by default, are logged and counted towards the key error limit. Modifying the ErrorConfiguration on the measure group or partition can control this.
ms345138.as2k5dataintegrity_05(en-US,SQL.90).gif
Figure 5. Edit Bindings dialog box
Note that the NullProcessing needs to be set on the KeyColumn of the measure group attribute. In the Dimension Usage tab of the cube designer, edit the relationship between the dimension and the measure group. Click Advanced, select the granularity attribute, and set the NullProcessing.

NULLs in Snowflaked Dimension Table

The product dimension table has records in which the product_class_id is NULL. This is handled in the same way as in the previous section, except that the NullProcessing needs to be set on the KeyColumn of DimensionAttribute (in the Properties pane of the dimension designer).

Inconsistent Relationships in Dimension Table

As described earlier, inconsistent relationships in the dimension table result in duplicate keys. In the example described earlier, the brand_name "Best Choice" appears twice with different product_class_id values. This produces a KeyDuplicate error that by default is ignored, and the server discards the duplicate record.
Alternatively setting KeyDuplicate=ReportAndContinue/ReportAndStop will cause the errors to be logged. The log can then be examined to determine potential flaws in the dimension design.

Conclusion

Data integrity issues can be challenging for database administrators to manage. SQL Server 2005 Analysis Services provides sophisticated controls such as Unknown Member, Null Processing, and Error Configuration that can greatly simplify cube management tasks.



Real Time rocessing

Real Time (Transaction)

Another type of real-time system involves transactions. As soon as a transaction is received by the computer, it is processed and any data files are updated. The system is real-time. If a customer books a seat at a performance then the details are added to the bookings file immediately and that seat is flagged as 'Reserved'. Another customer who asks a few seconds later for a seat at the same performance will not be able to book the same seat.

Example: Theater Booking System


Theater Booking System

About the Booking System

-  They built Theater booking systems to any specification
- Take payments for bookings by debit and credit cards online
- Theatre systems are web based and can be accessed from anywhere in the world
- Each system is custom designed to match your company branding
- Bookings can still be taken via phone and in person by using the administration panel
Real Time (Process)

Real-time means that data is processed immediately.

A real-time system is always 'up-to-date'. The computer used in a real-time system is 'dedicated' - it does nothing else.


An example of a real-time system is a Process-Control System - computers controlling a manufacturing system. Input data received from sensors is processed immediately, analyzed and any necessary actions taken without any delay.   


Another example would be a flight simulator.

KEY FIELD

 KEY FIELD

A key is a data item that allows us to uniquely identify individual occurrences or an entity type. You can sort and quickly retrieve information from a database by choosing one or more fields (ie attributes) to act as keys. For instance, in a student's table you could use a combination of the last name and first name fields (or perhaps last name, first name and birth dates to ensure you identify each student uniquely) as a key field. 
There are several types of key field:
  • Primary Key
  • Secondary Key
  • Foreign key
  • Simple key
  • Compound key
  • Composite key


Primary Key


A primary key consists of one or more attributes that distinguishes a specific record from any other. For each record in the table the primary key acts like a driver's licence number or a national insurance number, only one number exists for each person.
For example, your student number is a primary key as this uniquely identifies you within the college student records system. An employee number uniquely identifies a member of staff within a company. 
An IP address uniquely addresses a PC on the internet.
A primary key is mandatory. That is, each entity occurrence must have a value for its primary key.


Secondary Key

An entity may have one or more choices for the primary key. Collectively these are known as candidate keys. One is selected as the primary key. Those not selected are known as secondary keys.
For example, an employee has an employee number, a National Insurance (NI) number and an email address. If the employee number is chosen as the primary key then the NI number and email address are secondary keys. However, it is important to note that if any employee does not have a NI number or email address (ie: the attribute is not mandatory) then it cannot be chosen as a primary key.



Foreign Key
A foreign key is one or more attribute in one entity, which enables a link (or relationship) to another entity. That is, a foreign key in one entity links to a primary key in another entity. However, if the business rules permit, a foreign key may be optional.
For example, an employee works in a department. The department number column in the employee entity is a foreign key, which links to the department entity.
Foreign keys will be explained in more detail when we explore the normalisation process later in this section.

What is a Key field in a Database and How should I choose one?

Keys are crucial to a table structure for many reasons, some of which are identified below:
§                  They ensure that each record in a table is precisely identified.
§                  They help establish and enforce various types of integrity.
§                  They serve to establish table relationships.

Now let's see how you should choose your key(s). First, let's make up a little table to look at:
PersonID
LastName
FirstName
D.O.B
1
Smith
Robert
01/01/1970
2
Jones
Robert
01/01/1970
4
Smith
Henry
01/01/1970
5
Jones
Henry
01/01/1970

A superkey is a column or set of columns that uniquely identify a record. This table has many superkeys:
§                  PersonID
§                  PersonID + LastName
§                  PersonID + FirstName
§                  PersonID + DOB
§                  PersonID + LastName + FirstName
§                  PersonID + LastName + DOB
§                  PersonID + FirstName + DOB
§                  PersonID + LastName + FirstName + DOB
§                  LastName + FirstName + DOB

All of these will uniquely identify each record, so each one is a superkey. Of those keys, a key which is comprised of more than one column is a composite key; a key of only one column is a simple key.
A candidate key is a superkey that has no unique subset; it contains no columns that are not necessary to make it unique. 

This table has 2 candidate keys:
§                  PersonID
§                  LastName + FirstName + DOB

Not all candidate keys make good primary keys: Note that these may work for our current data set, but would likely be bad choices for future data. It is quite possible for two people to share a full name and date of birth.

We select a primary key from the candidate keys. This primary key will uniquely identify each record. It may or may not provide information about the record it identifies. It must not be Null-able, that is if it exists in a record it can not have the value Null. It must be unique. Itcan not be changed. Any candidate keys we do not select become alternate keys.

We will select (PersonID) as the primary key. This makes (LastName + FirstName + DOB) an alternate key.
Now, if this field PersonID is meaningful, that is it is used for any other purpose than making the record unique, it is a natural key or intelligent key. In this case PersonID is probablynot an AutoNumber field, but is rather a "customer number" for use, much like the UPC or ISBN.

However, if this field is not meaningful, that is it is strictly for the database to internally identify a unique record, it is a surrogate key or blind key. In this case Person ID probablyis an AutoNumber field, and it should not be used except internally by the database.

There is a long running debate over whether one should use natural or surrogate keys, and I'm not going to foolishly attempt to resolve it here. Whichever you use, stick with it. If you choose to generate an AutoNumber that is only used to identify a record, do not expose that number to the user. They will surely want to change it, and you can not change primary keys.

I can now use my chosen primary key in another table, to relate the two tables. It may or may not have the same name in that second table. In either case, with respect to the second table it is a foreign key, and if in that second table the foreign key field is not indexed it is a fast foreign key.
Many thanks to JasonM, mdbmakers.com Moderator, at www.MDBMAKERS.com for permission to use the above article