This is the corpus for
Hawes, Timothy (2009). COMPUTATIONAL ANALYSIS OF THE CONVERSATIONAL DYNAMICS OF THE UNITED STATES SUPREME COURT.

Each file is a cleaned version of the original case transcript PDF available at:
http://www.supremecourt.gov/oral_arguments/argument_transcripts.aspx

This corpus includes all timespans covered in the thesis.
Documents use a loose XML formatting. The structure is as follows:

   <DOC>
      <DOCNO>[CASE_ID]</DOCNO>
      <TAGS>
         <VOTE justice=[JUSTICE_ID] value=[VOTE_CODING]/>
         ...
         <VOTE .../>
      </TAGS>
      [BODY]
   </DOC>

Where:
	
CASE_ID is the Supreme Court ID given to the case

JUSTICE_ID is the last name of the justice whose vote is given by this tag (name internal punctuation removed)

VOTE_CODING specifies the votes of the indicated justice in this case in [PETITIONER_VOTE]::[RESPONDENT_VOTE] format, votes may be PRO or CON (e.g. PRO::CON was decided in favor of the Petitioner and against the Respondent). Some votes may be listed as NA or null. NA votes were explicitly coded as "did not participate" while "null" votes lacked data. Votes were obtained from the SPAETH Supreme Court database (http://scdb.wustl.edu/about.php). 

BODY includes one line per "transcript item". Transcript items include:
  Timestamps: such as "(10:02 a.m.)" or "(Whereupon, at 10:57 a.m., the case in the above-entitled matter was submitted.)"
  Case Section Headings: such as "REBUTTAL ARGUMENT OF LISA MADIGAN ON BEHALF OF THE PETITIONER" or "ORAL ARGUMENT OF RALPH E. MECZYK ON BEHALF OF THE RESPONDENT"
  Speaker Turns: such as "JUSTICE STEVENS: Thank you, Mr. Wray." or "MR. WRAY: That is correct, Justice Souter."
Transcript items are fairly consistent in their formatting, that is times appear in consistent locations enclosed in parenthesis or brackets, headings are usually all caps and contain keywords like "REBUTTAL", "ARGUMENT" or "ARGUMENTS" and turns reliably an all caps name followed by a colon then text. So, this means relatively straight forward regular expressions should be able to ID different segments. 

The following is a listing of the 179 cases used in the Classification Experiments and represent the second Roberts "Natural Court" as described in the thesis. 11 cases (again described in the thesis) were held out from this natural court due to coding inconsistencies:
04-1034
04-10566
04-1170b
04-1324
04-1327
04-1350
04-1360b
04-1376
04-1506
04-1527
04-1528
04-1544
04-1618
04-1704
04-1739
04-433
04-473b
04-607
04-9728
05-1056
05-1074
05-1120
05-1126
05-11284
05-11304
05-1157
05-1240
05-1256
05-1272
05-128
05-1284
05-130
05-1342
05-1345
05-1382
05-1429
05-1448
05-1508
05-1541
05-1575
05-1589
05-1629
05-1631
05-18
05-184
05-200
05-204
05-259
05-260
05-352
05-380
05-381
05-409
05-416
05-465
05-493
05-502
05-5224
05-547
05-5705
05-593
05-595
05-5966
05-5992
05-608
05-6551
05-669
05-705
05-7053
05-7058
05-746
05-785
05-83
05-848
05-85
05-8794
05-8820
05-908
05-915
05-9222
05-9264
05-983
05-996
05-998
06-1005
06-10119
06-102
06-1037
06-1082
06-11429
06-11543
06-116
06-11612
06-1164
06-1181
06-1195
06-1204
06-1221
06-1265
06-1286
06-1287
06-1321
06-1322
06-134
06-134orig
06-1413
06-1431
06-1456
06-1457
06-1463
06-1498
06-1505
06-1509
06-157
06-1646
06-1666
06-1717
06-179
06-219
06-278
06-313
06-340
06-376
06-413
06-427
06-43
06-457
06-480
06-484
06-5247
06-5306
06-531
06-5618
06-562
06-571
06-5754
06-593
06-618
06-6330
06-637
06-6407
06-666
06-6911
06-694
06-713
06-7517
06-766
06-7949
06-8120
06-8273
06-84
06-856
06-9130
06-923
06-937
06-939
06-969
06-984
06-989
07-208
07-21
07-210
07-214
07-219
07-290
07-308
07-312
07-320
07-330
07-343
07-371
07-411
07-440
07-455
07-474
07-5439
07-552
07-6053
07-77

