A short example to begin with
Nothing is better than an example to present the features of TraMineR.
We will use for this purpose a data set from
McVicar and Anyadike-Danes [2002]
which is freely downloadable from the internet.
The data was converted into an R data frame named
mvad that is distributed with the TraMineR package.
The
mvad data contains 72 monthly activity state variables from July 1993 to June 1999 for 712 individuals and a series of covariates.
We assume that the TraMineR library has been loaded with the
library(TraMineR) command.
Figure
1 shows examples of plots produced by TraMineR.
Commands used for generating the plots are given
below the Figure.
 a) Index plot of first 10 sequences |
 b) Index plot of all sequences sorted by gcse5eq |
 c) Frequency plot of 10 most frequent sequences |
 d) State distribution plot by time point |
 e) Legend |
 f) Histogram of sequence turbulence |
Figure 1: A short example using the mvad data
- Define a vector containing the legends for the states to appear in the graphics and create a sequence object which will be used as argument to the next functions
data(mvad)
mvad.labels <- c("employment", "further education", "higher education",
"joblessness", "school", "training")
mvad.scodes <- c("EM","FE","HE","JL","SC","TR")
mvad.seq <- seqdef(mvad, 15:86, states=mvad.scodes, labels=mvad.labels)
-
Plot of the 10 first sequences. Result in Fig. 1(a).
seqiplot(mvad.seq, withlegend=F)
-
Plot all sequences sorted by variable `gcse5eq', a binary dummy indicating qualifications gained by the end of compulsory education:
1 if 5 or more GCSEs at grades A-C or equivalent, 0 otherwise. Result in Fig. 1(b).
seqiplot(mvad.seq, tlim=0, sortv=mvad$gcse5eq, border=NA, space=0,
withlegend=F)
-
Draw the sequence frequency plot of the 10 most frequent sequences with bar width proportional to the frequencies. Result in Fig. 1(c).
seqfplot(mvad.seq, pbarw=T, withlegend=F)
-
Plot the state distribution by time points. Result in Fig. 1(d).
seqdplot(mvad.seq, withlegend=F)
-
Plot a single legend as a separate graphic since several plots use the same color codes for the states. Result in Fig. 1(e).
seqlegend(mvad.seq)
-
Compute, summarize and plot the histogram of the sequence turbulences. Result in Fig. 1(f).
mvad.turb <- seqST(mvad.seq)
summary(mvad.turb)
hist(mvad.turb, col="cyan")
-
Compute the optimal matching distances using substitution costs based on transition rates observed in the data and a 1 indel cost. The resulting distance matrix is stored in the dist.om1 object.
submat <- seqsubm(mvad.seq, method= "TRATE")
dist.om1 <- seqdist(mvad.seq, method="OM", indel=1, sm=submat)
- Make a typology of the trajectories: load the cluster
library, build a Ward hierarchical clustering of the sequences from the optimal matching distances and retrieve for each individual sequence the cluster membership of the 3 class solution. We do not show here the dendrogram produced by plot(clusterward1), which, indeed, is not a TraMineR feature.
library(cluster)
clusterward1 <- agnes(dist.om1, diss=TRUE, method="ward")
plot(clusterward1)
cl1.3 <- cutree(clusterward1, k=3)
cl1.3fac <- factor(cl1.3, labels = c("Type 1", "Type 2", "Type 3"))
- Plot the state distribution within each cluster. Result in Fig. 2.
seqdplot(mvad.seq, group=cl1.3fac)

Figure 2: Transversal state distributions within each cluster - mvad data
- Plot the 10 most frequent sequences of each cluster. Result in Fig. 3.
seqfplot(mvad.seq, group=cl1.3fac, pbarw=T)

Figure 3: The 10 most frequent sequences by cluster - mvad data
Instead of focusing on sequences of states, we can look at sequences of transitions or events.
TraMineR offers specific tools to deal with such kind of data. For dealing with such event sequences, we can:
- Create event sequences from the state sequences. The result is stored in an event sequence object.
mvad.seqe <- seqecreate(mvad.seq)
- Look for frequent event subsequences and plot the 15 most frequent ones. Result in Fig. 4.
fsubseq <- seqefsub(mvad.seqe, pMinSupport=0.05)
plot(fsubseq[1:15], col="cyan")

Figure 4: Frequencies of most frequent transitions - mvad data
- Determine the most discriminating transitions between clusters and plot the frequencies by cluster of the 6 first ones. Result in Fig. 5.
discr <- seqecmpgroup(fsubseq, group=cl1.3fac)
plot(discr[1:6])

Figure 5: Most discriminating transitions between clusters - mvad data
--------------
McVicar, D. and M. Anyadike-Danes (2002).
Predicting successful and unsuccessful transitions from school to
work by using sequence methods.
Journal of the Royal Statistical Society. Series A (Statistics
in Society) 165(2), 317-334.