Page 1¶
SPSS · Statistical Package for the Social Sciences¶
INTRODUCTION¶
The development of SPSS began in 1965 at Stanford University and through the years new facilities have continually been added. It is an integrated system of computer programs designed for the analysis of social science data. The system provides a unified and comprehensive package that enables the user to perform many different types of data analysis in a simple and convenient manner. SPSS allows a great deal of flexibility in the format of data. It provides the user with a comprehensive set of procedures for data transformation and file manipulation, and it offers the researcher a large number of statistical routines commonly used in the social sciences.
SPSS is a batch oriented system, but it may also be run interactively. The batch job is built up interactively with the aid of an editor.
FEATURES¶
SPSS performs over 40 major procedures including:
- Frequency distributions and bar charts.
- Descriptive statistics.
- Crosstabulation and measures of association.
- Descriptions of subpopulations.
- Tests for equality of two means.
- Scattergrams and correlations.
- Oneway and multiway analysis of variance.
- Nonparametric correlations and tests.
- Bivariate and multiple regression.
- Tabulation of multiple-response data.
- Partial and canonical correlation.
- Multivariate analysis of variance.
- Discriminant analysis.
- Factor analysis.
- Survival analysis.
- Box-Jenkins time series analysis.
- Report writing.
SPSS System is the only batch system that is truly user-oriented. It was invented and developed for researchers who have little or no computer experience. It obeys English language commands and gives easy-to-read, clearly labeled results.
A highly sophisticated tool – for statistical analysis, tabulation, report writing and general purpose data management.
Output may be customized to needs, because each procedure includes many options for handling and presenting data.
No limitations on the number of cases, and up to 5000 variables are allowed.
Designed to dramatically minimize costly CPU time.
The data-management facilities can be used to modify a file of data permanently and can also be used in conjunction with any of the statistical procedures. These facilities enable the user to generate new variables which are mathematical and/or logical combinations of existing variables, to recode variables, and to sample, select, or weight specified cases. Furthermore, the user can add or alter the data cases or the data-descriptioinal information in the file, such as labels, missing-value codes, etc.
SPSS is used in such applications as: - Survey and marketing data analysis. - Personnel studies. - Government report preparation. - Mailing list screening to increase response rates. - Census data analysis. - Peak load forecasting for utilities. - Statistical modeling of manufacturing processes. - Insurance claims and policyholder analysis. - Crime and drug abuse research. - Biological and medical research, including survival analysis. - Traffic pattern and accident report analysis.
A1–4000–0682
Page 2¶
The SPSS Job¶
All SPSS jobs require two files: the data file and the command file. The command file consists of SPSS commands entered into the computer as lines of input at a terminal in the same manner as data are entered. These commands declare the names of variables well known by, indicate where the information about a variable is located, specify whether to ignore a case if a variable has a certain value, assign a short label to a variable and to values for a variable, identify the type of statistical analysis to be performed, and tell SPSS when to start reading the data file. The commands must be entered in a sequence such that the computer always has the information needed to process the next command.
In the following respect, the files are permanent, but not immutable. The data, or any of the documenting information, may be added to, deleted, or altered at the user's will, and a new or updated file may be retained. Additional variables can be added to the file as well as additional cases; labels may be added or altered; new variables or scales can be created from existing ones; and documenting messages may be saved with the file. In short, the system file becomes a permanent self-documenting entity, and the user need only remember the name of the file and the order of the variables within. Even this information, if forgotten, can be retrieved easily.
The most important aspect of system files is the potential effect they can have (if properly used) on the interaction between researcher and data during day-to-day analysis. With a complicated data file, it will take considerable time to prepare and debug the initial run that defines the file. However, once this has been accomplished, massive runs taking a long time to plan and prepare need not and probably should not be made. With a system file the researcher can begin to explore particular themes and hypotheses, submitting frequent runs requiring little preparation, and thus the likelihood of errors is minimized.
Data entered into the SPSS system may be substructured into groups called subfiles. Subfiles may be sampling points such as cities; they may be national samples in crossnational surveys research; they may consist of data from different time trials or experimental treatments.
Once the subfile structure has been created, individual subfiles may be selected for processing, combinations of subfiles may be processed together, or the subfile structure may be ignored altogether.
A random sample of the cases in a file may be obtained, specific cases may be selected for processing, and the cases in the file may be weighted. The user is able to specify all the conditions and criteria for accomplishing sampling, selecting, and weighting during any processing run.
Research may involve dual levels of analysis or at least the examination of the impact of some larger unit or institution on the behavior of individuals. Subprogram AGGREGATE permits the researcher to define larger aggregation units and to compute aggregated variables.
SPSS Batch System collects data about its own use which exist on a special file. The Usage Data Collection Utility Program (UDC) reads this file:
The UDC-program can be run at any time for just reading the use of SPSS since last time the UDC-file was initiated. A report is then printed out. The UDC-file consists of only one record containing an identification code, the data of initialization, the last date SPSS was used and some counters showing the job size, the time spent and the frequency of various events.
The report which the UDC-program produces consists of 5 pages, and is divided into 6 main parts:
- Number of runs.
- CPU-time spent by the various statistics routines in SPSS.
- Control cards (i.e. number of records) used.
- Space used during the data transformation.
- Size of SPSS files.
- CPU-time spent by SPSS-runs.
Documentation¶
Norman H. Nie, C. Hadlai Hull, Jean G. Jenkins, Karin Steinbrenner and Dale H. Bent: SPSS-Statistical Package for the Social Sciences. McGraw – Hill Book Company.
______ ______
/ \ Norsk Data/ \
/ \ / \
/____/\____\________/____/\____\
\ NP PRODUKTER ______ /
\ Olav Helses vei 5 / /
\ Boks 5 / /
\ 1301 SANDVIKA / /
\______________/____/
______ ______
/ \ COMTEC / \
/ \ division / \
/____/\____\________/____/\____\
\ NP Jerikoveien 20 ______ /
\ Boks 4 Linderberg/ Parer /
\ Oslo 10 / /
\ Tel: 02-90030/ /
\______________/____/
(Note: ASCII representation of logos.)