Differences
This shows you the differences between two versions of the page.
| Next revision | Previous revision | ||
| eis:resources:annotation_report [2025/03/21 11:03] – created tom | eis:resources:annotation_report [2026/07/02 23:34] (current) – tom | ||
|---|---|---|---|
| Line 1: | Line 1: | ||
| ====== Report on Annotation strategy and outcomes ====== | ====== Report on Annotation strategy and outcomes ====== | ||
| + | |||
| + | //Research notes March 2025.// | ||
| + | |||
| ===== Method ===== | ===== Method ===== | ||
| + | <WRAP group> | ||
| + | <WRAP half column> | ||
| + | We construct a manual annotation task to determine stances towards instances of ethical issues in software related Reddit posts. | ||
| + | Two annotators, the author and a research assistant, annotate a dataset of 1000 posts. They follow this [[eis: | ||
| + | Annotators first conduct a trial annotation round to establish sufficient common understanding of the task. They annotate 100 randomly sampled posts from the full dataset of 65769 software related posts and compare their agreement before the main task. They must achieve a substantial agreement of Cohen' | ||
| + | |||
| + | The annotation uses a predefined set of labels, namely the ethical issue taxonomy established in [[https:// | ||
| + | The annotation guide instructs annotators to read the post and consider if the post contains an ethical issue in software. A post is annotated if it contains a description of an issue that fits the definition of an ethical issue type. The annotator selects all types that fit, having the option to define a new type in the rare case where none of the current types fit an apparent ethical issue. Annotators are instructed to highlight the piece of the text that led to their decision. | ||
| + | |||
| + | Lastly, annotators assign the stance label, determining the Redditor' | ||
| + | |||
| + | The two annotators are given a stratified dataset of 1000 posts. The annotator set two checkpoints, | ||
| + | Throughout the process, very complex cases are further to be discussed with researcher' | ||
| + | |||
| + | ==== Data ==== | ||
| + | |||
| + | After the trial sample was extracted and annotated, we decide to filter the dataset for posts with a higher likelihood to contain an ethical issue due to a very low ratio of ~20%. | ||
| + | From the 65769 posts, we extract a stratified sample of 4998 posts over all software domains, ensuring that at least one post from every subreddit is included. | ||
| + | We filter the posts by prompting Anthropic' | ||
| + | Claude is instructed to answer with TRUE or FALSE, and to resort to the former when in doubt. This increases the likelihood of not missing more complex cases. | ||
| + | We are aware of the limitations of this novel approach. We accept potential biases of Claude to ensure a higher density of interesting cases for analysis. This speeds up our approach to learn how to automatically assign stances towards ethical issues. Future work will establish approaches to filter occurrences of ethical issues more reliably. Works such as that of [[https:// | ||
| + | Claude assigned the TRUE label to 1756 posts, the FALSE label to 3227 posts and gibberish in 15 cases. | ||
| + | </ | ||
| + | <WRAP half column> | ||
| + | <wrap figure> | ||
| + | |||
| + | **Fig.1.** Number of posts per software domain. | ||
| + | </ | ||
| + | |||
| + | <wrap figure> | ||
| + | |||
| + | **Fig.2.** Claude labels. | ||
| + | </ | ||
| + | </ | ||
| + | </ | ||
| ===== Preparation ===== | ===== Preparation ===== | ||
| + | |||
| + | The annotators prepared by reading the annotation guidelines, reading the definitions of the ethical issue types and | ||
| ===== Preliminary results ===== | ===== Preliminary results ===== | ||
eis/resources/annotation_report.1742551385.txt.gz · Last modified: by tom