Reporting Experimental Results from SIR Objects
When conducting experiments using the objects contained in the repository, there are potential threats to statistical and conclusion validity that users should consider. Discussed here are various limitations of the resources provided by the repository and possible implications and ramifications of those limitations. Relevant threats and concerns should be addressed in publications reporting empirical results based upon the resources in this repository.
Quantity of Observations
In some cases, experimental results may yield a limited number of observations, a situation that limits the validity of statistical analysis and comparisons. This may sometimes be exacerbated by characteristics of the objects, such as a limited number of seeded faults, or provided test suites failing to reveal a significant number of faults. In these circumstances, there is a high risk that comparisons with published results, or even statistical comparisons between different treatments, are not valid. It is generally most appropriate to classify such experiments as limited case studies, and careful consideration should be given before classifying such an experiment as a controlled study. In a controlled study, these limitations represent a significant threat to internal validity and should be accordingly acknowledged.
One potential approach to address such limitations is to use
objects with a greater number of faults and larger test suites to
execute or augment experiments. These programs will facilitate
collection of a much larger set of observations, enabling stronger
and higher confidence statistical comparisons. Among the objects
currently available in the repository, this is an advantage of the
Siemens programs and space. However, these
particular objects are small (quite small in some cases), which
results in concerns about the ability to generalize experimental
results to programs of more "realistic" sizes. As
a consequence, this is a threat to external validity that also
needs to be acknowledged.
A strong approach to performing an empirical evaluation is often to perform controlled experiments on the smaller objects, permitting the study to consider causal effects with high internal validity, and then to perform case studies on larger objects to evaluate the generalization of results. A discussion of the tradeoff between internal and external validities is usually appropriate and can be used to justify higher confidence claims from such experiments.