On the efficacy of dynamic behavior comparison for judging functional equivalence
Kessel, Marcus
;
Atkinson, Colin
DOI:
|
https://doi.org/10.1109/SCAM.2019.00030
|
URL:
|
https://ieeexplore.ieee.org/document/8930875
|
Document Type:
|
Conference or workshop publication
|
Year of publication:
|
2019
|
Book title:
|
SCAM 2019 : 19th International Working Conference on Source Code Analysis and Manipulation, September 30 - October 1, 2019, Cleveland, Ohio : proceedings
|
Page range:
|
193-203
|
Conference title:
|
19th International Working Conference on Source Code Analysis and Manipulation (SCAM)
|
Location of the conference venue:
|
Cleveland, OH
|
Date of the conference:
|
30.09.-01.10.19
|
Publisher:
|
O'Conner, Lisa
|
Place of publication:
|
Los Alamitos, CA [u.a.]
|
Publishing house:
|
IEEE
|
ISBN:
|
978-1-7281-4938-7 , 978-1-7281-4937-0
|
ISSN:
|
1942-5430 , 2470-6892
|
Publication language:
|
English
|
Institution:
|
School of Business Informatics and Mathematics > Software Engineering (Atkinson 2003-)
|
Subject:
|
004 Computer science, internet
|
Abstract:
|
Since it was first proposed in 1992 under the name of "behavior sampling", the idea of judging whether software systems are functionally equivalent by observing their responses to common stimuli (i.e. tests) has been used for a range of tasks such as software retrieval, functional redundancy measurement and semantic clone detection. However, its efficacy has only been studied in one small experiment, with limited generalizability, described in the original paper proposing the approach. The results of that experiment suggest that a relatively small number of randomly generated tests (i.e. 4) is sufficient to recognize non-functional-equivalent software 85% of the time. This number has therefore been adopted as "sufficient" in numerous applications of the approach. In this paper we present a much larger study which suggests at least 39 randomly generated tests are actually needed to achieve this level of effectiveness, but that a far fewer number of tests generated using coverage-based heuristics are sufficient. Since these results are much more generalizable, they have implications for future applications of behavioral sampling for dynamic behavior comparison.
|
| Dieser Eintrag ist Teil der Universitätsbibliographie. |
Search Authors in
You have found an error? Please let us know about your desired correction here: E-Mail
Actions (login required)
|
Show item |
|
|