Wednesday, October 7, 2026

The Complete Research Paper Folder

 

How to Organize a Scientific Publication from First Draft to Final Publication

A scientific paper does not begin when you upload a manuscript to a journal, and it does not end when the paper receives a DOI.

Between the first draft and the final publication there may be:

  • dozens of manuscript versions,
  • multiple sets of figures,
  • statistical analyses,
  • raw and processed datasets,
  • code,
  • plagiarism checks,
  • AI-use checks,
  • internal reviews,
  • preprint versions,
  • several journal submissions,
  • reviewer reports,
  • responses to reviewers,
  • revised manuscripts,
  • rejected versions,
  • accepted manuscripts,
  • proofs,
  • copyright agreements,
  • APC waiver applications,
  • repository deposits,
  • social-media posts,
  • press releases,
  • blog posts,
  • and many other documents.

If these materials are stored in an ad hoc collection of files such as

paper_final.docx
paper_final2.docx
paper_final_revised.docx
paper_FINAL.docx
paper_FINAL_really_final.docx
paper_FINAL_really_final2.docx

then the problem is not simply untidy file management.

It becomes a research reproducibility and provenance problem.

A better approach is to create a standardized Publication Master Folder for every paper.

The principle is simple:

Every important scientific, editorial, administrative and public-facing stage of a paper should have a predictable location and an identifiable version.


1. The publication folder should represent the entire life of the paper

A useful top-level structure is:

PAPER_PROJECT/
│
├── 00_PROJECT_ADMIN/
├── 01_RESEARCH_DATA/
├── 02_ANALYSIS/
├── 03_CODE/
├── 04_RESULTS/
├── 05_FIGURES/
├── 06_TABLES/
├── 07_MANUSCRIPT_DEVELOPMENT/
├── 08_INTEGRITY_CHECKS/
├── 09_INTERNAL_REVIEW/
├── 10_PREPRINT/
├── 11_JOURNAL_SUBMISSIONS/
├── 12_PEER_REVIEW/
├── 13_REVISIONS/
├── 14_ACCEPTANCE/
├── 15_PRODUCTION/
├── 16_PUBLICATION/
├── 17_APC/
├── 18_DATA_CODE_REPOSITORIES/
├── 19_PUBLIC_COMMUNICATION/
├── 20_ARCHIVE/
└── README.md

This is deliberately more extensive than what a journal requires.

The journal may receive only a fraction of these files.

The rest form the permanent provenance record of the paper.


2. Start with a README file

The root directory should contain:

README.md

This is the first file someone should read when entering the project.

It should contain:

Paper title:
Short title:
Project ID:
Corresponding author:
First author:
Lab:
Institution:

Current status:
Current manuscript version:
Current journal:
Submission date:
Revision status:
Preprint DOI:
Journal DOI:

Repository:
Code repository:
Data repository:

Last updated:
Maintained by:

It should also contain a short history:

2026-01-10  Project initiated
2026-03-15  First complete manuscript
2026-04-02  Internal lab review completed
2026-04-10  bioRxiv v1 posted
2026-04-20  Submitted to Journal A
2026-06-15  Rejected
2026-06-25  Submitted to Journal B
2026-08-03  Major revision
2026-09-01  Revised manuscript submitted
2026-09-25  Accepted
2026-10-05  Published

This simple file can become extraordinarily valuable years later.


3. 00_PROJECT_ADMIN

This folder contains the administrative information about the paper.

00_PROJECT_ADMIN/
│
├── Project_Overview/
├── Author_List/
├── Author_Contributions/
├── Affiliations/
├── ORCID/
├── Funding/
├── Grant_Information/
├── Conflict_of_Interest/
├── Ethics/
├── Permissions/
├── Institutional_Approvals/
├── Data_Management/
├── Publication_Agreements/
└── Timeline/

For example:

Author_List/
├── Author_List_v01.xlsx
├── Author_List_v02.xlsx
└── Author_Order_Final.pdf

The final author order should be explicitly archived.

This avoids future confusion about whether an author was added, removed or moved during manuscript development.


4. 01_RESEARCH_DATA

This is the scientific foundation of the paper.

01_RESEARCH_DATA/
│
├── 01_RAW_DATA/
├── 02_METADATA/
├── 03_PROCESSED_DATA/
├── 04_FINAL_ANALYSIS_DATA/
├── 05_SOURCE_DATA_FOR_FIGURES/
├── 06_SOURCE_DATA_FOR_TABLES/
├── 07_EXTERNAL_DATA/
├── 08_DATABASE_DOWNLOADS/
├── 09_SEQUENCE_DATA/
├── 10_METADATA_DICTIONARIES/
└── 11_DATA_RELEASE/

The distinction between raw, processed, and analysis-ready data is extremely important.

Never overwrite raw data.

For example:

01_RAW_DATA/
02_PROCESSED_DATA/
03_FINAL_ANALYSIS_DATA/

should represent a direction of processing:

RAW → PROCESSED → ANALYSIS

not three interchangeable copies of the same files.


5. 02_ANALYSIS

This contains the actual scientific analyses.

For computational biology:

02_ANALYSIS/
│
├── 01_Data_Cleaning/
├── 02_Quality_Control/
├── 03_Sequence_Analysis/
├── 04_Alignment/
├── 05_Phylogenetics/
├── 06_Gene_Loss/
├── 07_Statistics/
├── 08_Robustness_Analysis/
├── 09_Sensitivity_Analysis/
├── 10_Alternative_Models/
└── 11_Final_Analysis/

This is where one should preserve analyses that may not ultimately appear in the paper.

That is important.

An analysis that was abandoned can explain why the final analysis looks the way it does.


6. 03_CODE

Code deserves its own dedicated directory.

03_CODE/
│
├── R/
├── Python/
├── Shell/
├── Workflow/
├── Configuration/
├── Environment/
├── Notebooks/
├── Containers/
├── Tests/
└── README.md

For example:

03_CODE/R/
├── 01_import_data.R
├── 02_clean_data.R
├── 03_statistics.R
├── 04_PCA.R
├── 05_phylogeny.R
└── 06_make_figures.R

Also preserve computational environments:

Environment/
├── environment.yml
├── requirements.txt
├── renv.lock
└── software_versions.txt

For serious computational projects, the goal should be:

Someone should be able to determine what software, packages and versions generated the published results.


7. 04_RESULTS

This folder contains the outputs of analyses.

04_RESULTS/
│
├── 01_Raw_Analysis_Output/
├── 02_Statistics/
├── 03_Model_Output/
├── 04_Phylogenetic_Trees/
├── 05_Alignments/
├── 06_Tables/
├── 07_Figure_Source_Data/
└── 08_Final_Results/

Do not mix raw computational output with polished publication figures.


8. 05_FIGURES

This deserves particularly careful organization.

05_FIGURES/
│
├── 01_Working_Figures/
├── 02_Manuscript_Figures/
├── 03_Supplementary_Figures/
├── 04_Editable_Source/
├── 05_High_Resolution/
├── 06_Web_Resolution/
├── 07_Graphical_Abstract/
└── 08_Figure_Source_Data/

For example:

Figure_01/
├── Figure_01_editable.ai
├── Figure_01_final.pdf
├── Figure_01_300dpi.tif
├── Figure_01_web.png
└── Figure_01_source_data.xlsx

The source data underlying a figure should not disappear after publication.


9. 06_TABLES

06_TABLES/
│
├── Working/
├── Main_Text/
├── Supplementary/
├── Source_Data/
└── Publication_Final/

Tables should preferably remain in editable formats.


10. 07_MANUSCRIPT_DEVELOPMENT

This is where the manuscript evolves before it is submitted to a journal.

07_MANUSCRIPT_DEVELOPMENT/
│
├── 01_Outline/
├── 02_First_Draft/
├── 03_Internal_Drafts/
├── 04_Complete_Drafts/
├── 05_Author_Revisions/
├── 06_Tracked_Changes/
├── 07_Clean_Copies/
├── 08_Final_PreSubmission/
└── 09_Archived_Drafts/

Do not rely solely on filenames such as final.docx.

Use explicit versions:

Manuscript_v01.docx
Manuscript_v02.docx
Manuscript_v03.docx

Better still, associate versions with dates:

2026-05-10_Manuscript_v03.docx

A consistent naming convention makes files easier to retrieve and interpret later.


11. 08_INTEGRITY_CHECKS

This is a particularly important folder.

08_INTEGRITY_CHECKS/
│
├── 01_Plagiarism/
├── 02_AI_Usage/
├── 03_Image_Integrity/
├── 04_Data_Integrity/
├── 05_Reference_Check/
├── 06_Fact_Check/
├── 07_Statistical_Check/
├── 08_Code_Check/
├── 09_Authorship_Check/
└── 10_Final_Integrity_Clearance/

Plagiarism

01_Plagiarism/
├── Originality_Report_v01.pdf
├── Originality_Report_v02.pdf
├── Similarity_Report.pdf
├── Exclusions_Used.txt
└── Author_Review_of_Matches.pdf

Do not merely save the similarity percentage.

Save the actual report.

Also record:

  • software used
  • date
  • manuscript version
  • similarity score
  • exclusions applied
  • interpretation of flagged sections

A 15% similarity score without context is much less informative than the underlying report.


12. AI-usage documentation

Because AI-assisted writing and analysis are increasingly part of research workflows, create:

02_AI_Usage/
│
├── AI_Usage_Log.xlsx
├── AI_Disclosure_for_Journal.txt
├── AI_Check_Report.pdf
├── AI_Generated_Text_Reviewed/
├── AI_Assisted_Code/
├── AI_Assisted_Analysis/
├── AI_Image_Use/
└── Final_AI_Compliance_Check.pdf

The AI usage log might contain:

DateToolPurposeMaterialHuman verification
2026-05-02AI toolGrammarIntroductionYes
2026-05-03AI toolCode debuggingR scriptYes
2026-05-04AI toolLiterature organizationReferencesYes

The purpose is not to create unnecessary bureaucracy.

It is to preserve an auditable record of how AI contributed to the research process.


13. Image-integrity checks

For papers containing microscopy, western blots, gels or other image-based data:

03_Image_Integrity/
├── Original_Images/
├── Processed_Images/
├── Image_Processing_Log.xlsx
├── Figure_Comparison/
├── Integrity_Check_Report.pdf
└── Final_Approval.pdf

Never let the only surviving copy of an experimental image be the version embedded in PowerPoint or Word.


14. Statistical and data-integrity checks

Create a final verification folder:

04_Data_Integrity/
├── Raw_vs_Reported_Values.xlsx
├── Figure_Source_Verification.xlsx
├── Table_Source_Verification.xlsx
└── Final_Data_Check.pdf

This can answer:

Does every number in the manuscript trace back to an identifiable analysis or source dataset?


15. 09_INTERNAL_REVIEW

Before submission, have an independent internal review.

09_INTERNAL_REVIEW/
│
├── 01_Scientific_Review/
├── 02_Methodological_Review/
├── 03_Statistical_Review/
├── 04_Language_Review/
├── 05_Figure_Review/
├── 06_Data_Review/
├── 07_Senior_Author_Review/
├── 08_Final_Checklist/
└── 09_Approval/

Keep comments and responses.

For example:

Internal_Reviewer_01_comments.docx
Internal_Reviewer_01_response.docx

This creates a useful history of scientific decision-making.


16. 10_PREPRINT

If a preprint is appropriate, create a dedicated folder.

10_PREPRINT/
│
├── 00_Preprint_Decision/
├── 01_Preprint_Manuscript/
├── 02_Preprint_Figures/
├── 03_Preprint_Supplement/
├── 04_Submission_Package/
├── 05_Submission_Record/
├── 06_bioRxiv/
├── 07_arXiv/
├── 08_Preprint_Versions/
├── 09_Preprint_DOI/
└── 10_Preprint_to_Journal_Mapping/

The last folder is particularly useful.

Record:

bioRxiv v1 → Journal manuscript v1
bioRxiv v2 → Journal manuscript v3
Journal accepted manuscript → Published article

bioRxiv allows revised versions before formal acceptance, and versions retain the same basic DOI while version-specific URLs identify particular versions.

Therefore, do not treat a preprint as a disposable PDF.

It is part of the publication history.


17. Preprint file-size management

Create:

10_PREPRINT/
└── 04_Submission_Package/
    ├── Full_Resolution/
    ├── Compressed/
    ├── PDF/
    ├── Source/
    └── Upload_Archive/

For example:

Preprint_v01_full.zip
Preprint_v01_submission.zip
Preprint_v01.pdf
Preprint_v01_source.zip

Keep the exact package that was uploaded.

If figures had to be compressed to meet an upload limit, retain both:

Figure_01_original.tif
Figure_01_preprint_compressed.jpg

Do not overwrite the original.


18. 11_JOURNAL_SUBMISSIONS

This is one of the most important folders.

Never assume that there is one manuscript.

There may be five different journal-specific versions.

Use:

11_JOURNAL_SUBMISSIONS/
│
├── Journal_A/
├── Journal_B/
├── Journal_C/
└── Journal_D/

Each journal gets its own complete package.

For example:

Journal_A/
│
├── 00_Journal_Guidelines/
├── 01_Submission/
├── 02_Cover_Letter/
├── 03_Figures/
├── 04_Tables/
├── 05_Supplement/
├── 06_Declarations/
├── 07_Reviewer_Suggestions/
├── 08_Submission_Portal/
├── 09_Submission_Confirmation/
├── 10_Editorial_Correspondence/
└── 11_Submitted_Package/

This prevents a common disaster:

submitting a manuscript formatted for Journal A to Journal B while accidentally retaining Journal A's declarations, references or supplementary numbering.


19. Preserve the exact submitted package

Inside every journal folder:

Submitted_Package/
├── Manuscript_SUBMITTED.pdf
├── Manuscript_SUBMITTED.docx
├── Figure_01_SUBMITTED.tif
├── Figure_02_SUBMITTED.tif
├── Supplement_SUBMITTED.pdf
├── Cover_Letter_SUBMITTED.pdf
└── Submission_Metadata.pdf

The word SUBMITTED is important.

This is the immutable historical record.

Never replace these files with later versions.


20. 12_PEER_REVIEW

Once the paper enters peer review:

12_PEER_REVIEW/
│
├── Journal_A/
│   ├── Round_1/
│   │   ├── Reviewer_1.pdf
│   │   ├── Reviewer_2.pdf
│   │   ├── Reviewer_3.pdf
│   │   ├── Editor_Letter.pdf
│   │   └── Decision_Letter.pdf
│   │
│   └── Round_2/
│
└── Journal_B/

Preserve the original reviewer reports.


21. 13_REVISIONS

This folder should capture the entire revision history.

13_REVISIONS/
│
├── Journal_A/
│   ├── Round_1_Major_Revision/
│   ├── Round_2_Minor_Revision/
│   └── Final_Acceptance/
│
└── Journal_B/

Each round should contain:

Round_1/
├── Decision_Letter.pdf
├── Reviewer_Comments/
├── Response_to_Reviewers/
├── Revised_Manuscript_Tracked.docx
├── Revised_Manuscript_Clean.docx
├── Revised_Figures/
├── Revised_Supplement/
└── Submitted_Revision_Package/

22. The response-to-reviewers deserves its own history

For example:

Response_to_Reviewers/
├── Response_v01.docx
├── Response_v02.docx
├── Response_FINAL.docx
└── Response_SUBMITTED.pdf

The final submitted response should never be overwritten.


23. 14_ACCEPTANCE

Once accepted:

14_ACCEPTANCE/
│
├── Acceptance_Letter/
├── Accepted_Manuscript/
├── Final_Figures/
├── Final_Supplement/
├── Final_Data/
├── Final_Code/
├── Copyright/
├── License/
├── Author_Agreement/
└── Publication_Charges/

This is the transition from peer-reviewed manuscript to publication production.


24. 15_PRODUCTION

The production stage often generates a completely new set of files.

15_PRODUCTION/
│
├── Typesetting/
├── Proofs/
├── Proof_Corrections/
├── Author_Corrections/
├── Production_Correspondence/
├── Final_Proof/
├── XML_or_JATS/
├── Published_PDF/
└── Version_of_Record/

Save every proof.

For example:

Proof_v01.pdf
Proof_v02.pdf
Proof_FINAL.pdf

If a production error occurs years later, these records can be extremely useful.


25. 16_PUBLICATION

Once the article is published:

16_PUBLICATION/
│
├── DOI/
├── Published_PDF/
├── HTML/
├── Supplementary/
├── Published_Figures/
├── Published_Tables/
├── Article_Metadata/
├── Citation/
├── Publisher_Page/
└── Final_Bibliographic_Record/

Record:

  • DOI
  • publication date
  • volume
  • issue
  • article number/pages
  • publisher
  • journal
  • PMID
  • PubMed Central ID, if applicable
  • Crossref information
  • repository links

26. 17_APC

APC management deserves a separate folder.

17_APC/
│
├── 01_APC_Eligibility/
├── 02_Waiver_Request/
├── 03_Supporting_Documents/
├── 04_Waiver_Correspondence/
├── 05_Waiver_Decision/
├── 06_Discount/
├── 07_Invoice/
├── 08_Payment/
└── 09_Final_APC_Record/

For example:

02_Waiver_Request/
├── APC_Waiver_Letter.docx
├── Institutional_Proof.pdf
├── Funding_Statement.pdf
└── Supporting_Explanation.pdf

Then:

04_Waiver_Correspondence/
├── Publisher_Request.pdf
├── Author_Response.pdf
└── Final_Decision.pdf

The financial history of the publication is therefore preserved separately from the scientific record.


27. 18_DATA_CODE_REPOSITORIES

The paper may have multiple public repositories.

18_DATA_CODE_REPOSITORIES/
│
├── GitHub/
├── GitLab/
├── Zenodo/
├── OSF/
├── Dryad/
├── Figshare/
├── GenBank/
├── SRA/
├── GEO/
├── ENA/
├── TreeBASE/
└── Other_Databases/

Maintain:

Repository_Register.xlsx

with:

RepositoryMaterialAccession/DOIVersionDate
ZenodoCodeDOIv1date
SRARaw readsaccession—date
GitHubCodeURLrelease 1.0date

Repository versioning is important because later changes should not silently alter the exact computational object supporting a published result. Persistent versioning systems such as Zenodo explicitly preserve separate versions and identifiers.


28. 19_PUBLIC_COMMUNICATION

Publication should not be the end.

Create:

19_PUBLIC_COMMUNICATION/
│
├── 01_Plain_Language_Summary/
├── 02_Blog_Post/
├── 03_Press_Release/
├── 04_Graphical_Abstract/
├── 05_Social_Media/
├── 06_Lab_Website/
├── 07_Institutional_Website/
├── 08_Presentation_Slides/
├── 09_Video/
├── 10_Podcast/
├── 11_FAQ/
└── 12_Media_Enquiries/

29. Blog post folder

For example:

02_Blog_Post/
├── Blog_Draft_v01.docx
├── Blog_Draft_v02.docx
├── Blog_Final.docx
├── Images/
├── References/
├── Published/
└── URL.txt

The blog post should ideally be prepared before publication, so that it can be released soon after the paper appears.

The blog should link to:

  • published article
  • DOI
  • preprint
  • data
  • code
  • supplementary information

30. Social-media folder

05_Social_Media/
│
├── X/
├── LinkedIn/
├── Instagram/
├── ResearchGate/
├── Bluesky/
└── Institutional/

Within each:

Post_v01.txt
Post_Final.txt
Image.png

Prepare several lengths:

01_One_Line.txt
02_Short.txt
03_Medium.txt
04_Long.txt

31. 20_ARCHIVE

This is the final preservation layer.

After publication, create:

20_ARCHIVE/
│
├── 01_FINAL_MANUSCRIPT/
├── 02_FINAL_DATA/
├── 03_FINAL_CODE/
├── 04_FINAL_FIGURES/
├── 05_FINAL_TABLES/
├── 06_FINAL_SUPPLEMENT/
├── 07_SUBMISSION_HISTORY/
├── 08_REVIEW_HISTORY/
├── 09_PUBLICATION_HISTORY/
├── 10_REPOSITORY_RECORD/
├── 11_CORRESPONDENCE/
├── 12_INTEGRITY_RECORDS/
├── 13_APC_RECORDS/
└── 14_README/

The archive should contain enough information to reconstruct the publication history without requiring someone to search through an email inbox.


32. Use a publication version numbering system

A useful convention is:

v01
v02
v03
...

But version numbers should have meaning.

For example:

M01 = manuscript development version
P01 = preprint version
J1.1 = Journal 1 first submission
J1.R1 = Journal 1 revision 1
J1.R2 = Journal 1 revision 2
J2.1 = Journal 2 first submission
ACC = accepted manuscript
VOR = version of record

Then filenames become:

Manuscript_M03.docx
Manuscript_bioRxiv_v01.pdf
Manuscript_J1.1_SUBMITTED.pdf
Manuscript_J1.R1_SUBMITTED.pdf
Manuscript_J1.R2_SUBMITTED.pdf
Manuscript_J1_ACC.pdf
Manuscript_VOR.pdf

This is much safer than:

final.docx
final2.docx
final_new.docx
final_latest.docx

33. Never overwrite historical versions

This is perhaps the single most important rule.

If:

Manuscript_J1.1_SUBMITTED.pdf

was actually submitted, never modify it.

If the manuscript changes, create:

Manuscript_J1.R1_SUBMITTED.pdf

The same principle applies to:

  • figures
  • supplementary files
  • code
  • datasets
  • response letters
  • cover letters
  • repository releases

Versioning is not unnecessary duplication: it establishes provenance. Research-data versioning guidance emphasizes the importance of being able to identify exactly which version underlies a published result.


34. Keep a MASTER CHANGE LOG

At the root:

CHANGELOG.md

For example:

# CHANGELOG

## 2026-04-01 — v01
Initial complete manuscript.

## 2026-04-10 — v02
Added phylogenetic analysis.
Revised Figure 3.
Updated Discussion.

## 2026-04-20 — bioRxiv v1
Preprint deposited.

## 2026-05-15 — Journal A submission
Major formatting changes.
Added Supplementary Figure S4.

## 2026-07-01 — Revision 1
Addressed Reviewer 1 comments.
Reanalysed dataset.
Figure 2 replaced.

## 2026-08-15 — Accepted
Final accepted manuscript archived.

## 2026-09-01 — Published
DOI assigned.

This may ultimately be more valuable than a complicated file-management system.


35. Keep a submission register

Create:

SUBMISSION_REGISTER.xlsx

with:

JournalVersionSubmittedDecisionReasonNext journal
Journal AJ1.110 AprRejectScopeJournal B
Journal BJ2.115 MayReviseMajor revision—
Journal BJ2.R120 JulAccept——

This prevents a paper's history from becoming dependent on someone's memory.


36. Keep a manuscript–figure–data mapping

A particularly powerful addition is:

FIGURE_DATA_MAP.xlsx

For every figure:

FigureSource dataAnalysis scriptRaw dataFinal file
Fig 1data_01.csvscript_01.Rraw_01Fig1.tif
Fig 2data_03.csvscript_05.Rraw_03Fig2.tif

Do the same for tables.

This creates a direct chain:

RAW DATA
   ↓
ANALYSIS
   ↓
SOURCE DATA
   ↓
FIGURE/TABLE
   ↓
MANUSCRIPT
   ↓
PUBLISHED ARTICLE

That is the essence of reproducible publication.


37. Keep correspondence

Create:

CORRESPONDENCE/
│
├── Authors/
├── Journal/
├── Editor/
├── Reviewers/
├── Publisher/
├── Repository/
├── Institution/
└── Funding_APC/

Important email correspondence should be exported or otherwise archived where institutional policy permits.

Do not rely on a personal inbox as the only record.


38. The final archive should be self-contained

The ultimate goal is to create something like:

ARTICLE_ARCHIVE/
│
├── README.md
├── CHANGELOG.md
├── PUBLICATION_METADATA.txt
├── SUBMISSION_REGISTER.xlsx
├── FIGURE_DATA_MAP.xlsx
├── MANUSCRIPT/
├── DATA/
├── CODE/
├── FIGURES/
├── TABLES/
├── SUPPLEMENT/
├── INTEGRITY/
├── PREPRINT/
├── JOURNAL_SUBMISSIONS/
├── PEER_REVIEW/
├── REVISIONS/
├── ACCEPTANCE/
├── PRODUCTION/
├── APC/
├── REPOSITORIES/
├── PUBLIC_COMMUNICATION/
└── CORRESPONDENCE/

A compressed archive of the appropriate release can then be preserved as a long-term snapshot.

Repositories such as Zenodo can preserve uploaded collections and provide persistent identifiers; when a collection contains multiple files/folders, a compressed archive can also be used where appropriate.


39. But don't put everything into one ZIP file

There is an important distinction between:

Working storage

and

archival storage.

Your working directory may contain:

temporary/
scratch/
old/
test/
debug/

These should not automatically become part of the public archive.

Instead, create a deliberate release package:

RELEASE/
├── README.md
├── DATA/
├── CODE/
├── FIGURES/
├── TABLES/
├── SUPPLEMENT/
├── LICENSE
└── CITATION.cff

Only materials necessary for understanding, reproducing or reusing the published work should enter the public release.


40. A simple rule for deciding where a file belongs

Every file should answer one of these questions:

A. Is this scientific evidence?

Put it in:

DATA/
ANALYSIS/
RESULTS/

B. Is this how the evidence was generated?

Put it in:

CODE/
WORKFLOW/
ENVIRONMENT/

C. Is this a manuscript version?

Put it in:

MANUSCRIPT/
SUBMISSIONS/
REVISIONS/

D. Is this evidence of research integrity?

Put it in:

INTEGRITY/

E. Is this evidence of communication with a journal?

Put it in:

SUBMISSIONS/
PEER_REVIEW/
CORRESPONDENCE/

F. Is this evidence of publication?

Put it in:

ACCEPTANCE/
PRODUCTION/
PUBLICATION/

G. Is this about communicating the research to the public?

Put it in:

PUBLIC_COMMUNICATION/

41. The golden rules

A lab-wide publication-management system can be reduced to about fifteen rules:

1. Never call a file simply final.

2. Never overwrite a submitted version.

3. Never overwrite raw data.

4. Preserve the exact files actually submitted to every journal.

5. Keep rejected submissions.

A rejected manuscript is part of the publication history.

6. Keep reviewer reports and responses.

7. Keep plagiarism/similarity reports.

8. Keep AI-use documentation and declarations.

9. Keep original figure/source data.

10. Keep code and software environments.

11. Record repository accession numbers and DOI versions.

12. Separate journal-specific submission packages.

13. Keep APC waiver requests and decisions.

14. Prepare communication material before publication rather than after publication.

15. Create a final immutable archive once the paper is published.


42. The ultimate purpose

The purpose of such a folder structure is not bureaucracy.

It is scientific provenance.

A well-organized paper should allow someone to move backwards through the entire history:

Published article
       ↓
Version of record
       ↓
Accepted manuscript
       ↓
Final revision
       ↓
Reviewer comments
       ↓
Original submission
       ↓
Preprint
       ↓
Manuscript development
       ↓
Figures and tables
       ↓
Analysis
       ↓
Code
       ↓
Processed data
       ↓
Raw data

And it should also allow someone to move forward:

Published article
       ↓
DOI
       ↓
Data repository
       ↓
Code repository
       ↓
Preprint
       ↓
Blog post
       ↓
Institutional news
       ↓
Public communication

That means the paper becomes more than a PDF.

It becomes a traceable research object.


43. The ideal lab philosophy

A useful philosophy for a research group is:

A publication should be reproducible, auditable, traceable and recoverable.

Reproducible means another researcher can understand how the results were produced.

Auditable means there is evidence for important scientific and editorial decisions.

Traceable means every important result can be connected to its source data and analysis.

Recoverable means that years later, the lab can reconstruct what was submitted, revised, accepted and published.

This approach also aligns naturally with FAIR-oriented research-data management, which emphasizes findability, accessibility, interoperability and reusability.


44. A final recommended master structure

If I were implementing this as a standard for a computational biology laboratory, I would use:

PAPER_ID_SHORT_TITLE/
│
├── 00_ADMIN/
│   ├── Authors/
│   ├── Funding/
│   ├── Ethics/
│   ├── Permissions/
│   ├── Timeline/
│   └── README.md
│
├── 01_DATA/
│   ├── Raw/
│   ├── Metadata/
│   ├── Processed/
│   ├── Analysis_Ready/
│   └── Source_Data/
│
├── 02_ANALYSIS/
│   ├── QC/
│   ├── Primary/
│   ├── Secondary/
│   ├── Robustness/
│   └── Sensitivity/
│
├── 03_CODE/
│   ├── R/
│   ├── Python/
│   ├── Shell/
│   ├── Workflow/
│   ├── Environment/
│   └── README.md
│
├── 04_RESULTS/
│
├── 05_FIGURES/
│   ├── Working/
│   ├── Main/
│   ├── Supplementary/
│   ├── Editable/
│   └── Source_Data/
│
├── 06_TABLES/
│
├── 07_MANUSCRIPT/
│   ├── Drafts/
│   ├── Tracked_Changes/
│   ├── Clean/
│   └── PreSubmission/
│
├── 08_INTEGRITY/
│   ├── Plagiarism/
│   ├── AI_Usage/
│   ├── Image_Integrity/
│   ├── Data_Integrity/
│   ├── Statistics/
│   └── References/
│
├── 09_INTERNAL_REVIEW/
│
├── 10_PREPRINT/
│   ├── bioRxiv/
│   ├── arXiv/
│   ├── Versions/
│   ├── Submission/
│   └── DOI/
│
├── 11_JOURNAL_SUBMISSIONS/
│   ├── Journal_A/
│   ├── Journal_B/
│   ├── Journal_C/
│   └── Journal_D/
│
├── 12_PEER_REVIEW/
│
├── 13_REVISIONS/
│   ├── Round_1/
│   ├── Round_2/
│   └── Final/
│
├── 14_ACCEPTANCE/
│
├── 15_PRODUCTION/
│   ├── Proofs/
│   ├── Corrections/
│   └── Final/
│
├── 16_PUBLICATION/
│
├── 17_APC/
│   ├── Eligibility/
│   ├── Waiver/
│   ├── Supporting_Documents/
│   ├── Correspondence/
│   └── Payment/
│
├── 18_REPOSITORIES/
│   ├── Code/
│   ├── Data/
│   ├── Sequences/
│   ├── Preprint/
│   └── Metadata/
│
├── 19_PUBLIC_COMMUNICATION/
│   ├── Blog/
│   ├── Press_Release/
│   ├── Social_Media/
│   ├── Graphical_Abstract/
│   ├── Website/
│   └── Presentations/
│
├── 20_CORRESPONDENCE/
│
├── 21_FINAL_ARCHIVE/
│
├── README.md
├── CHANGELOG.md
├── SUBMISSION_REGISTER.xlsx
└── FIGURE_DATA_MAP.xlsx

The important conceptual change

The biggest change I would make compared with the way most researchers organize papers is this:

Don't create a folder called "Paper."

Create a folder called the paper's entire lifecycle.

The manuscript is only one component of that lifecycle.

The same folder should tell the story of the research from:

raw data → analysis → manuscript → integrity checks → preprint → journal submission → peer review → revision → acceptance → production → publication → repository → public communication → long-term archive.

That is what turns ordinary file storage into a research information management system.


No comments: