How to Organize a Scientific Publication from First Draft to Final Publication
A scientific paper does not begin when you upload a manuscript to a journal, and it does not end when the paper receives a DOI.
Between the first draft and the final publication there may be:
- dozens of manuscript versions,
- multiple sets of figures,
- statistical analyses,
- raw and processed datasets,
- code,
- plagiarism checks,
- AI-use checks,
- internal reviews,
- preprint versions,
- several journal submissions,
- reviewer reports,
- responses to reviewers,
- revised manuscripts,
- rejected versions,
- accepted manuscripts,
- proofs,
- copyright agreements,
- APC waiver applications,
- repository deposits,
- social-media posts,
- press releases,
- blog posts,
- and many other documents.
If these materials are stored in an ad hoc collection of files such as
paper_final.docx
paper_final2.docx
paper_final_revised.docx
paper_FINAL.docx
paper_FINAL_really_final.docx
paper_FINAL_really_final2.docxthen the problem is not simply untidy file management.
It becomes a research reproducibility and provenance problem.
A better approach is to create a standardized Publication Master Folder for every paper.
The principle is simple:
Every important scientific, editorial, administrative and public-facing stage of a paper should have a predictable location and an identifiable version.
1. The publication folder should represent the entire life of the paper
A useful top-level structure is:
PAPER_PROJECT/
│
├── 00_PROJECT_ADMIN/
├── 01_RESEARCH_DATA/
├── 02_ANALYSIS/
├── 03_CODE/
├── 04_RESULTS/
├── 05_FIGURES/
├── 06_TABLES/
├── 07_MANUSCRIPT_DEVELOPMENT/
├── 08_INTEGRITY_CHECKS/
├── 09_INTERNAL_REVIEW/
├── 10_PREPRINT/
├── 11_JOURNAL_SUBMISSIONS/
├── 12_PEER_REVIEW/
├── 13_REVISIONS/
├── 14_ACCEPTANCE/
├── 15_PRODUCTION/
├── 16_PUBLICATION/
├── 17_APC/
├── 18_DATA_CODE_REPOSITORIES/
├── 19_PUBLIC_COMMUNICATION/
├── 20_ARCHIVE/
└── README.mdThis is deliberately more extensive than what a journal requires.
The journal may receive only a fraction of these files.
The rest form the permanent provenance record of the paper.
2. Start with a README file
The root directory should contain:
README.mdThis is the first file someone should read when entering the project.
It should contain:
Paper title:
Short title:
Project ID:
Corresponding author:
First author:
Lab:
Institution:
Current status:
Current manuscript version:
Current journal:
Submission date:
Revision status:
Preprint DOI:
Journal DOI:
Repository:
Code repository:
Data repository:
Last updated:
Maintained by:It should also contain a short history:
2026-01-10 Project initiated
2026-03-15 First complete manuscript
2026-04-02 Internal lab review completed
2026-04-10 bioRxiv v1 posted
2026-04-20 Submitted to Journal A
2026-06-15 Rejected
2026-06-25 Submitted to Journal B
2026-08-03 Major revision
2026-09-01 Revised manuscript submitted
2026-09-25 Accepted
2026-10-05 PublishedThis simple file can become extraordinarily valuable years later.
3. 00_PROJECT_ADMIN
This folder contains the administrative information about the paper.
00_PROJECT_ADMIN/
│
├── Project_Overview/
├── Author_List/
├── Author_Contributions/
├── Affiliations/
├── ORCID/
├── Funding/
├── Grant_Information/
├── Conflict_of_Interest/
├── Ethics/
├── Permissions/
├── Institutional_Approvals/
├── Data_Management/
├── Publication_Agreements/
└── Timeline/For example:
Author_List/
├── Author_List_v01.xlsx
├── Author_List_v02.xlsx
└── Author_Order_Final.pdfThe final author order should be explicitly archived.
This avoids future confusion about whether an author was added, removed or moved during manuscript development.
4. 01_RESEARCH_DATA
This is the scientific foundation of the paper.
01_RESEARCH_DATA/
│
├── 01_RAW_DATA/
├── 02_METADATA/
├── 03_PROCESSED_DATA/
├── 04_FINAL_ANALYSIS_DATA/
├── 05_SOURCE_DATA_FOR_FIGURES/
├── 06_SOURCE_DATA_FOR_TABLES/
├── 07_EXTERNAL_DATA/
├── 08_DATABASE_DOWNLOADS/
├── 09_SEQUENCE_DATA/
├── 10_METADATA_DICTIONARIES/
└── 11_DATA_RELEASE/The distinction between raw, processed, and analysis-ready data is extremely important.
Never overwrite raw data.
For example:
01_RAW_DATA/
02_PROCESSED_DATA/
03_FINAL_ANALYSIS_DATA/should represent a direction of processing:
RAW → PROCESSED → ANALYSISnot three interchangeable copies of the same files.
5. 02_ANALYSIS
This contains the actual scientific analyses.
For computational biology:
02_ANALYSIS/
│
├── 01_Data_Cleaning/
├── 02_Quality_Control/
├── 03_Sequence_Analysis/
├── 04_Alignment/
├── 05_Phylogenetics/
├── 06_Gene_Loss/
├── 07_Statistics/
├── 08_Robustness_Analysis/
├── 09_Sensitivity_Analysis/
├── 10_Alternative_Models/
└── 11_Final_Analysis/This is where one should preserve analyses that may not ultimately appear in the paper.
That is important.
An analysis that was abandoned can explain why the final analysis looks the way it does.
6. 03_CODE
Code deserves its own dedicated directory.
03_CODE/
│
├── R/
├── Python/
├── Shell/
├── Workflow/
├── Configuration/
├── Environment/
├── Notebooks/
├── Containers/
├── Tests/
└── README.mdFor example:
03_CODE/R/
├── 01_import_data.R
├── 02_clean_data.R
├── 03_statistics.R
├── 04_PCA.R
├── 05_phylogeny.R
└── 06_make_figures.RAlso preserve computational environments:
Environment/
├── environment.yml
├── requirements.txt
├── renv.lock
└── software_versions.txtFor serious computational projects, the goal should be:
Someone should be able to determine what software, packages and versions generated the published results.
7. 04_RESULTS
This folder contains the outputs of analyses.
04_RESULTS/
│
├── 01_Raw_Analysis_Output/
├── 02_Statistics/
├── 03_Model_Output/
├── 04_Phylogenetic_Trees/
├── 05_Alignments/
├── 06_Tables/
├── 07_Figure_Source_Data/
└── 08_Final_Results/Do not mix raw computational output with polished publication figures.
8. 05_FIGURES
This deserves particularly careful organization.
05_FIGURES/
│
├── 01_Working_Figures/
├── 02_Manuscript_Figures/
├── 03_Supplementary_Figures/
├── 04_Editable_Source/
├── 05_High_Resolution/
├── 06_Web_Resolution/
├── 07_Graphical_Abstract/
└── 08_Figure_Source_Data/For example:
Figure_01/
├── Figure_01_editable.ai
├── Figure_01_final.pdf
├── Figure_01_300dpi.tif
├── Figure_01_web.png
└── Figure_01_source_data.xlsxThe source data underlying a figure should not disappear after publication.
9. 06_TABLES
06_TABLES/
│
├── Working/
├── Main_Text/
├── Supplementary/
├── Source_Data/
└── Publication_Final/Tables should preferably remain in editable formats.
10. 07_MANUSCRIPT_DEVELOPMENT
This is where the manuscript evolves before it is submitted to a journal.
07_MANUSCRIPT_DEVELOPMENT/
│
├── 01_Outline/
├── 02_First_Draft/
├── 03_Internal_Drafts/
├── 04_Complete_Drafts/
├── 05_Author_Revisions/
├── 06_Tracked_Changes/
├── 07_Clean_Copies/
├── 08_Final_PreSubmission/
└── 09_Archived_Drafts/Do not rely solely on filenames such as final.docx.
Use explicit versions:
Manuscript_v01.docx
Manuscript_v02.docx
Manuscript_v03.docxBetter still, associate versions with dates:
2026-05-10_Manuscript_v03.docxA consistent naming convention makes files easier to retrieve and interpret later.
11. 08_INTEGRITY_CHECKS
This is a particularly important folder.
08_INTEGRITY_CHECKS/
│
├── 01_Plagiarism/
├── 02_AI_Usage/
├── 03_Image_Integrity/
├── 04_Data_Integrity/
├── 05_Reference_Check/
├── 06_Fact_Check/
├── 07_Statistical_Check/
├── 08_Code_Check/
├── 09_Authorship_Check/
└── 10_Final_Integrity_Clearance/Plagiarism
01_Plagiarism/
├── Originality_Report_v01.pdf
├── Originality_Report_v02.pdf
├── Similarity_Report.pdf
├── Exclusions_Used.txt
└── Author_Review_of_Matches.pdfDo not merely save the similarity percentage.
Save the actual report.
Also record:
- software used
- date
- manuscript version
- similarity score
- exclusions applied
- interpretation of flagged sections
A 15% similarity score without context is much less informative than the underlying report.
12. AI-usage documentation
Because AI-assisted writing and analysis are increasingly part of research workflows, create:
02_AI_Usage/
│
├── AI_Usage_Log.xlsx
├── AI_Disclosure_for_Journal.txt
├── AI_Check_Report.pdf
├── AI_Generated_Text_Reviewed/
├── AI_Assisted_Code/
├── AI_Assisted_Analysis/
├── AI_Image_Use/
└── Final_AI_Compliance_Check.pdfThe AI usage log might contain:
| Date | Tool | Purpose | Material | Human verification |
|---|---|---|---|---|
| 2026-05-02 | AI tool | Grammar | Introduction | Yes |
| 2026-05-03 | AI tool | Code debugging | R script | Yes |
| 2026-05-04 | AI tool | Literature organization | References | Yes |
The purpose is not to create unnecessary bureaucracy.
It is to preserve an auditable record of how AI contributed to the research process.
13. Image-integrity checks
For papers containing microscopy, western blots, gels or other image-based data:
03_Image_Integrity/
├── Original_Images/
├── Processed_Images/
├── Image_Processing_Log.xlsx
├── Figure_Comparison/
├── Integrity_Check_Report.pdf
└── Final_Approval.pdfNever let the only surviving copy of an experimental image be the version embedded in PowerPoint or Word.
14. Statistical and data-integrity checks
Create a final verification folder:
04_Data_Integrity/
├── Raw_vs_Reported_Values.xlsx
├── Figure_Source_Verification.xlsx
├── Table_Source_Verification.xlsx
└── Final_Data_Check.pdfThis can answer:
Does every number in the manuscript trace back to an identifiable analysis or source dataset?
15. 09_INTERNAL_REVIEW
Before submission, have an independent internal review.
09_INTERNAL_REVIEW/
│
├── 01_Scientific_Review/
├── 02_Methodological_Review/
├── 03_Statistical_Review/
├── 04_Language_Review/
├── 05_Figure_Review/
├── 06_Data_Review/
├── 07_Senior_Author_Review/
├── 08_Final_Checklist/
└── 09_Approval/Keep comments and responses.
For example:
Internal_Reviewer_01_comments.docx
Internal_Reviewer_01_response.docxThis creates a useful history of scientific decision-making.
16. 10_PREPRINT
If a preprint is appropriate, create a dedicated folder.
10_PREPRINT/
│
├── 00_Preprint_Decision/
├── 01_Preprint_Manuscript/
├── 02_Preprint_Figures/
├── 03_Preprint_Supplement/
├── 04_Submission_Package/
├── 05_Submission_Record/
├── 06_bioRxiv/
├── 07_arXiv/
├── 08_Preprint_Versions/
├── 09_Preprint_DOI/
└── 10_Preprint_to_Journal_Mapping/The last folder is particularly useful.
Record:
bioRxiv v1 → Journal manuscript v1
bioRxiv v2 → Journal manuscript v3
Journal accepted manuscript → Published articlebioRxiv allows revised versions before formal acceptance, and versions retain the same basic DOI while version-specific URLs identify particular versions.
Therefore, do not treat a preprint as a disposable PDF.
It is part of the publication history.
17. Preprint file-size management
Create:
10_PREPRINT/
└── 04_Submission_Package/
├── Full_Resolution/
├── Compressed/
├── PDF/
├── Source/
└── Upload_Archive/For example:
Preprint_v01_full.zip
Preprint_v01_submission.zip
Preprint_v01.pdf
Preprint_v01_source.zipKeep the exact package that was uploaded.
If figures had to be compressed to meet an upload limit, retain both:
Figure_01_original.tif
Figure_01_preprint_compressed.jpgDo not overwrite the original.
18. 11_JOURNAL_SUBMISSIONS
This is one of the most important folders.
Never assume that there is one manuscript.
There may be five different journal-specific versions.
Use:
11_JOURNAL_SUBMISSIONS/
│
├── Journal_A/
├── Journal_B/
├── Journal_C/
└── Journal_D/Each journal gets its own complete package.
For example:
Journal_A/
│
├── 00_Journal_Guidelines/
├── 01_Submission/
├── 02_Cover_Letter/
├── 03_Figures/
├── 04_Tables/
├── 05_Supplement/
├── 06_Declarations/
├── 07_Reviewer_Suggestions/
├── 08_Submission_Portal/
├── 09_Submission_Confirmation/
├── 10_Editorial_Correspondence/
└── 11_Submitted_Package/This prevents a common disaster:
submitting a manuscript formatted for Journal A to Journal B while accidentally retaining Journal A's declarations, references or supplementary numbering.
19. Preserve the exact submitted package
Inside every journal folder:
Submitted_Package/
├── Manuscript_SUBMITTED.pdf
├── Manuscript_SUBMITTED.docx
├── Figure_01_SUBMITTED.tif
├── Figure_02_SUBMITTED.tif
├── Supplement_SUBMITTED.pdf
├── Cover_Letter_SUBMITTED.pdf
└── Submission_Metadata.pdfThe word SUBMITTED is important.
This is the immutable historical record.
Never replace these files with later versions.
20. 12_PEER_REVIEW
Once the paper enters peer review:
12_PEER_REVIEW/
│
├── Journal_A/
│ ├── Round_1/
│ │ ├── Reviewer_1.pdf
│ │ ├── Reviewer_2.pdf
│ │ ├── Reviewer_3.pdf
│ │ ├── Editor_Letter.pdf
│ │ └── Decision_Letter.pdf
│ │
│ └── Round_2/
│
└── Journal_B/Preserve the original reviewer reports.
21. 13_REVISIONS
This folder should capture the entire revision history.
13_REVISIONS/
│
├── Journal_A/
│ ├── Round_1_Major_Revision/
│ ├── Round_2_Minor_Revision/
│ └── Final_Acceptance/
│
└── Journal_B/Each round should contain:
Round_1/
├── Decision_Letter.pdf
├── Reviewer_Comments/
├── Response_to_Reviewers/
├── Revised_Manuscript_Tracked.docx
├── Revised_Manuscript_Clean.docx
├── Revised_Figures/
├── Revised_Supplement/
└── Submitted_Revision_Package/22. The response-to-reviewers deserves its own history
For example:
Response_to_Reviewers/
├── Response_v01.docx
├── Response_v02.docx
├── Response_FINAL.docx
└── Response_SUBMITTED.pdfThe final submitted response should never be overwritten.
23. 14_ACCEPTANCE
Once accepted:
14_ACCEPTANCE/
│
├── Acceptance_Letter/
├── Accepted_Manuscript/
├── Final_Figures/
├── Final_Supplement/
├── Final_Data/
├── Final_Code/
├── Copyright/
├── License/
├── Author_Agreement/
└── Publication_Charges/This is the transition from peer-reviewed manuscript to publication production.
24. 15_PRODUCTION
The production stage often generates a completely new set of files.
15_PRODUCTION/
│
├── Typesetting/
├── Proofs/
├── Proof_Corrections/
├── Author_Corrections/
├── Production_Correspondence/
├── Final_Proof/
├── XML_or_JATS/
├── Published_PDF/
└── Version_of_Record/Save every proof.
For example:
Proof_v01.pdf
Proof_v02.pdf
Proof_FINAL.pdfIf a production error occurs years later, these records can be extremely useful.
25. 16_PUBLICATION
Once the article is published:
16_PUBLICATION/
│
├── DOI/
├── Published_PDF/
├── HTML/
├── Supplementary/
├── Published_Figures/
├── Published_Tables/
├── Article_Metadata/
├── Citation/
├── Publisher_Page/
└── Final_Bibliographic_Record/Record:
- DOI
- publication date
- volume
- issue
- article number/pages
- publisher
- journal
- PMID
- PubMed Central ID, if applicable
- Crossref information
- repository links
26. 17_APC
APC management deserves a separate folder.
17_APC/
│
├── 01_APC_Eligibility/
├── 02_Waiver_Request/
├── 03_Supporting_Documents/
├── 04_Waiver_Correspondence/
├── 05_Waiver_Decision/
├── 06_Discount/
├── 07_Invoice/
├── 08_Payment/
└── 09_Final_APC_Record/For example:
02_Waiver_Request/
├── APC_Waiver_Letter.docx
├── Institutional_Proof.pdf
├── Funding_Statement.pdf
└── Supporting_Explanation.pdfThen:
04_Waiver_Correspondence/
├── Publisher_Request.pdf
├── Author_Response.pdf
└── Final_Decision.pdfThe financial history of the publication is therefore preserved separately from the scientific record.
27. 18_DATA_CODE_REPOSITORIES
The paper may have multiple public repositories.
18_DATA_CODE_REPOSITORIES/
│
├── GitHub/
├── GitLab/
├── Zenodo/
├── OSF/
├── Dryad/
├── Figshare/
├── GenBank/
├── SRA/
├── GEO/
├── ENA/
├── TreeBASE/
└── Other_Databases/Maintain:
Repository_Register.xlsxwith:
| Repository | Material | Accession/DOI | Version | Date |
|---|---|---|---|---|
| Zenodo | Code | DOI | v1 | date |
| SRA | Raw reads | accession | — | date |
| GitHub | Code | URL | release 1.0 | date |
Repository versioning is important because later changes should not silently alter the exact computational object supporting a published result. Persistent versioning systems such as Zenodo explicitly preserve separate versions and identifiers.
28. 19_PUBLIC_COMMUNICATION
Publication should not be the end.
Create:
19_PUBLIC_COMMUNICATION/
│
├── 01_Plain_Language_Summary/
├── 02_Blog_Post/
├── 03_Press_Release/
├── 04_Graphical_Abstract/
├── 05_Social_Media/
├── 06_Lab_Website/
├── 07_Institutional_Website/
├── 08_Presentation_Slides/
├── 09_Video/
├── 10_Podcast/
├── 11_FAQ/
└── 12_Media_Enquiries/29. Blog post folder
For example:
02_Blog_Post/
├── Blog_Draft_v01.docx
├── Blog_Draft_v02.docx
├── Blog_Final.docx
├── Images/
├── References/
├── Published/
└── URL.txtThe blog post should ideally be prepared before publication, so that it can be released soon after the paper appears.
The blog should link to:
- published article
- DOI
- preprint
- data
- code
- supplementary information
30. Social-media folder
05_Social_Media/
│
├── X/
├── LinkedIn/
├── Instagram/
├── ResearchGate/
├── Bluesky/
└── Institutional/Within each:
Post_v01.txt
Post_Final.txt
Image.pngPrepare several lengths:
01_One_Line.txt
02_Short.txt
03_Medium.txt
04_Long.txt31. 20_ARCHIVE
This is the final preservation layer.
After publication, create:
20_ARCHIVE/
│
├── 01_FINAL_MANUSCRIPT/
├── 02_FINAL_DATA/
├── 03_FINAL_CODE/
├── 04_FINAL_FIGURES/
├── 05_FINAL_TABLES/
├── 06_FINAL_SUPPLEMENT/
├── 07_SUBMISSION_HISTORY/
├── 08_REVIEW_HISTORY/
├── 09_PUBLICATION_HISTORY/
├── 10_REPOSITORY_RECORD/
├── 11_CORRESPONDENCE/
├── 12_INTEGRITY_RECORDS/
├── 13_APC_RECORDS/
└── 14_README/The archive should contain enough information to reconstruct the publication history without requiring someone to search through an email inbox.
32. Use a publication version numbering system
A useful convention is:
v01
v02
v03
...But version numbers should have meaning.
For example:
M01 = manuscript development version
P01 = preprint version
J1.1 = Journal 1 first submission
J1.R1 = Journal 1 revision 1
J1.R2 = Journal 1 revision 2
J2.1 = Journal 2 first submission
ACC = accepted manuscript
VOR = version of recordThen filenames become:
Manuscript_M03.docx
Manuscript_bioRxiv_v01.pdf
Manuscript_J1.1_SUBMITTED.pdf
Manuscript_J1.R1_SUBMITTED.pdf
Manuscript_J1.R2_SUBMITTED.pdf
Manuscript_J1_ACC.pdf
Manuscript_VOR.pdfThis is much safer than:
final.docx
final2.docx
final_new.docx
final_latest.docx33. Never overwrite historical versions
This is perhaps the single most important rule.
If:
Manuscript_J1.1_SUBMITTED.pdfwas actually submitted, never modify it.
If the manuscript changes, create:
Manuscript_J1.R1_SUBMITTED.pdfThe same principle applies to:
- figures
- supplementary files
- code
- datasets
- response letters
- cover letters
- repository releases
Versioning is not unnecessary duplication: it establishes provenance. Research-data versioning guidance emphasizes the importance of being able to identify exactly which version underlies a published result.
34. Keep a MASTER CHANGE LOG
At the root:
CHANGELOG.mdFor example:
# CHANGELOG
## 2026-04-01 — v01
Initial complete manuscript.
## 2026-04-10 — v02
Added phylogenetic analysis.
Revised Figure 3.
Updated Discussion.
## 2026-04-20 — bioRxiv v1
Preprint deposited.
## 2026-05-15 — Journal A submission
Major formatting changes.
Added Supplementary Figure S4.
## 2026-07-01 — Revision 1
Addressed Reviewer 1 comments.
Reanalysed dataset.
Figure 2 replaced.
## 2026-08-15 — Accepted
Final accepted manuscript archived.
## 2026-09-01 — Published
DOI assigned.This may ultimately be more valuable than a complicated file-management system.
35. Keep a submission register
Create:
SUBMISSION_REGISTER.xlsxwith:
| Journal | Version | Submitted | Decision | Reason | Next journal |
|---|---|---|---|---|---|
| Journal A | J1.1 | 10 Apr | Reject | Scope | Journal B |
| Journal B | J2.1 | 15 May | Revise | Major revision | — |
| Journal B | J2.R1 | 20 Jul | Accept | — | — |
This prevents a paper's history from becoming dependent on someone's memory.
36. Keep a manuscript–figure–data mapping
A particularly powerful addition is:
FIGURE_DATA_MAP.xlsxFor every figure:
| Figure | Source data | Analysis script | Raw data | Final file |
|---|---|---|---|---|
| Fig 1 | data_01.csv | script_01.R | raw_01 | Fig1.tif |
| Fig 2 | data_03.csv | script_05.R | raw_03 | Fig2.tif |
Do the same for tables.
This creates a direct chain:
RAW DATA
↓
ANALYSIS
↓
SOURCE DATA
↓
FIGURE/TABLE
↓
MANUSCRIPT
↓
PUBLISHED ARTICLEThat is the essence of reproducible publication.
37. Keep correspondence
Create:
CORRESPONDENCE/
│
├── Authors/
├── Journal/
├── Editor/
├── Reviewers/
├── Publisher/
├── Repository/
├── Institution/
└── Funding_APC/Important email correspondence should be exported or otherwise archived where institutional policy permits.
Do not rely on a personal inbox as the only record.
38. The final archive should be self-contained
The ultimate goal is to create something like:
ARTICLE_ARCHIVE/
│
├── README.md
├── CHANGELOG.md
├── PUBLICATION_METADATA.txt
├── SUBMISSION_REGISTER.xlsx
├── FIGURE_DATA_MAP.xlsx
├── MANUSCRIPT/
├── DATA/
├── CODE/
├── FIGURES/
├── TABLES/
├── SUPPLEMENT/
├── INTEGRITY/
├── PREPRINT/
├── JOURNAL_SUBMISSIONS/
├── PEER_REVIEW/
├── REVISIONS/
├── ACCEPTANCE/
├── PRODUCTION/
├── APC/
├── REPOSITORIES/
├── PUBLIC_COMMUNICATION/
└── CORRESPONDENCE/A compressed archive of the appropriate release can then be preserved as a long-term snapshot.
Repositories such as Zenodo can preserve uploaded collections and provide persistent identifiers; when a collection contains multiple files/folders, a compressed archive can also be used where appropriate.
39. But don't put everything into one ZIP file
There is an important distinction between:
Working storage
and
archival storage.
Your working directory may contain:
temporary/
scratch/
old/
test/
debug/These should not automatically become part of the public archive.
Instead, create a deliberate release package:
RELEASE/
├── README.md
├── DATA/
├── CODE/
├── FIGURES/
├── TABLES/
├── SUPPLEMENT/
├── LICENSE
└── CITATION.cffOnly materials necessary for understanding, reproducing or reusing the published work should enter the public release.
40. A simple rule for deciding where a file belongs
Every file should answer one of these questions:
A. Is this scientific evidence?
Put it in:
DATA/
ANALYSIS/
RESULTS/B. Is this how the evidence was generated?
Put it in:
CODE/
WORKFLOW/
ENVIRONMENT/C. Is this a manuscript version?
Put it in:
MANUSCRIPT/
SUBMISSIONS/
REVISIONS/D. Is this evidence of research integrity?
Put it in:
INTEGRITY/E. Is this evidence of communication with a journal?
Put it in:
SUBMISSIONS/
PEER_REVIEW/
CORRESPONDENCE/F. Is this evidence of publication?
Put it in:
ACCEPTANCE/
PRODUCTION/
PUBLICATION/G. Is this about communicating the research to the public?
Put it in:
PUBLIC_COMMUNICATION/41. The golden rules
A lab-wide publication-management system can be reduced to about fifteen rules:
1. Never call a file simply final.
2. Never overwrite a submitted version.
3. Never overwrite raw data.
4. Preserve the exact files actually submitted to every journal.
5. Keep rejected submissions.
A rejected manuscript is part of the publication history.
6. Keep reviewer reports and responses.
7. Keep plagiarism/similarity reports.
8. Keep AI-use documentation and declarations.
9. Keep original figure/source data.
10. Keep code and software environments.
11. Record repository accession numbers and DOI versions.
12. Separate journal-specific submission packages.
13. Keep APC waiver requests and decisions.
14. Prepare communication material before publication rather than after publication.
15. Create a final immutable archive once the paper is published.
42. The ultimate purpose
The purpose of such a folder structure is not bureaucracy.
It is scientific provenance.
A well-organized paper should allow someone to move backwards through the entire history:
Published article
↓
Version of record
↓
Accepted manuscript
↓
Final revision
↓
Reviewer comments
↓
Original submission
↓
Preprint
↓
Manuscript development
↓
Figures and tables
↓
Analysis
↓
Code
↓
Processed data
↓
Raw dataAnd it should also allow someone to move forward:
Published article
↓
DOI
↓
Data repository
↓
Code repository
↓
Preprint
↓
Blog post
↓
Institutional news
↓
Public communicationThat means the paper becomes more than a PDF.
It becomes a traceable research object.
43. The ideal lab philosophy
A useful philosophy for a research group is:
A publication should be reproducible, auditable, traceable and recoverable.
Reproducible means another researcher can understand how the results were produced.
Auditable means there is evidence for important scientific and editorial decisions.
Traceable means every important result can be connected to its source data and analysis.
Recoverable means that years later, the lab can reconstruct what was submitted, revised, accepted and published.
This approach also aligns naturally with FAIR-oriented research-data management, which emphasizes findability, accessibility, interoperability and reusability.
44. A final recommended master structure
If I were implementing this as a standard for a computational biology laboratory, I would use:
PAPER_ID_SHORT_TITLE/
│
├── 00_ADMIN/
│ ├── Authors/
│ ├── Funding/
│ ├── Ethics/
│ ├── Permissions/
│ ├── Timeline/
│ └── README.md
│
├── 01_DATA/
│ ├── Raw/
│ ├── Metadata/
│ ├── Processed/
│ ├── Analysis_Ready/
│ └── Source_Data/
│
├── 02_ANALYSIS/
│ ├── QC/
│ ├── Primary/
│ ├── Secondary/
│ ├── Robustness/
│ └── Sensitivity/
│
├── 03_CODE/
│ ├── R/
│ ├── Python/
│ ├── Shell/
│ ├── Workflow/
│ ├── Environment/
│ └── README.md
│
├── 04_RESULTS/
│
├── 05_FIGURES/
│ ├── Working/
│ ├── Main/
│ ├── Supplementary/
│ ├── Editable/
│ └── Source_Data/
│
├── 06_TABLES/
│
├── 07_MANUSCRIPT/
│ ├── Drafts/
│ ├── Tracked_Changes/
│ ├── Clean/
│ └── PreSubmission/
│
├── 08_INTEGRITY/
│ ├── Plagiarism/
│ ├── AI_Usage/
│ ├── Image_Integrity/
│ ├── Data_Integrity/
│ ├── Statistics/
│ └── References/
│
├── 09_INTERNAL_REVIEW/
│
├── 10_PREPRINT/
│ ├── bioRxiv/
│ ├── arXiv/
│ ├── Versions/
│ ├── Submission/
│ └── DOI/
│
├── 11_JOURNAL_SUBMISSIONS/
│ ├── Journal_A/
│ ├── Journal_B/
│ ├── Journal_C/
│ └── Journal_D/
│
├── 12_PEER_REVIEW/
│
├── 13_REVISIONS/
│ ├── Round_1/
│ ├── Round_2/
│ └── Final/
│
├── 14_ACCEPTANCE/
│
├── 15_PRODUCTION/
│ ├── Proofs/
│ ├── Corrections/
│ └── Final/
│
├── 16_PUBLICATION/
│
├── 17_APC/
│ ├── Eligibility/
│ ├── Waiver/
│ ├── Supporting_Documents/
│ ├── Correspondence/
│ └── Payment/
│
├── 18_REPOSITORIES/
│ ├── Code/
│ ├── Data/
│ ├── Sequences/
│ ├── Preprint/
│ └── Metadata/
│
├── 19_PUBLIC_COMMUNICATION/
│ ├── Blog/
│ ├── Press_Release/
│ ├── Social_Media/
│ ├── Graphical_Abstract/
│ ├── Website/
│ └── Presentations/
│
├── 20_CORRESPONDENCE/
│
├── 21_FINAL_ARCHIVE/
│
├── README.md
├── CHANGELOG.md
├── SUBMISSION_REGISTER.xlsx
└── FIGURE_DATA_MAP.xlsxThe important conceptual change
The biggest change I would make compared with the way most researchers organize papers is this:
Don't create a folder called "Paper."
Create a folder called the paper's entire lifecycle.
The manuscript is only one component of that lifecycle.
The same folder should tell the story of the research from:
raw data → analysis → manuscript → integrity checks → preprint → journal submission → peer review → revision → acceptance → production → publication → repository → public communication → long-term archive.
That is what turns ordinary file storage into a research information management system.
No comments:
Post a Comment