Versions Compared

Key

  • This line was added.
  • This line was removed.
  • Formatting was changed.

...

Develop and integrate internal instance of classifier code - We should integrate the classifier code into the arXiv production system rather than using API to code running on Paul Ginsparg's research machine. This was agreed by the SAB on 2013-09. Work was postponed in summer 2014 to allow quick initial deployment and to allow Paul Ginsparg time to tidy his code. There are uncertainties here because we haven't seen Paul's code and perhaps when we do we will want to rewrite some of the client-side code to reflect that understanding.

Queued for 2016: put a link to list to include:

 Better protect submitter email addresses - Currently submitter email addresses are stored in the metadata (.abs) files and in the listings files. While the web user interface does hide this information, the email addresses to too easily harvestable and potentially misused. We have had complaints from users where collaborating services have made the emails available. We should instead keep the submitter email addresses only in the database on the Cornell servers, and not in the metadata or listings files.

Assign DOIs to data - We accept data as ancillary files http://arxiv.org/help/ancillary_files but offer relatively little support. It would be more helpful to assign DataCite DOIs from EZID to ancillary files thus making them citeable.

 

Add store of "first processed" versions of arXiv articles in order to store an archival version of the processed copy the submitter saw and be able to understand any possible changes due to later reprocessing - The majority of arXiv submissions are made as TeX/LaTeX source files which arXiv then processes to produce PDF and other readable formats. While we take every care to maintain our TeX system in a way that will reliably reprocess old submissions, we currently have no facility to keep a copy of the first processed version. Such a store, accessible only to admins in the first instance, should be added in parallel to the usual cache of processed versions (/cache/ps_cache). This store should be populated by a a script which ensure that an existing processed version for a particular article-version is never overwritten by a new one.

 

User Support and Moderation

...