1. Controller and Contact
This Privacy Policy describes how user information is handled in Humanitext Antiqua (the “Service”).
- Controller: the Humanitext project (Center for Digital Humanities and Social Sciences, Nagoya University).
- Address: Nagoya University, Furo-cho, Chikusa-ku, Nagoya 464-8601, Japan.
- Requests for access, correction, or deletion, and complaints: the contact form, under the procedure in Section 7.
2. Information We Collect
- Account information: when you sign in with a Google account, your account identifier (UID), email address, and display name. Authentication is performed through Google Firebase Authentication, and the authentication state (your sign-in session) is handled by Firebase.
- Usage statistics: the corpus scope selected, the response mode, processing time, approximate token counts, success or failure, and language setting. These records contain no part of your questions or the answers, no account identifier, no email address, and no IP address.
- Conversation records: your questions, the generated answers, and the sources cited. These are recorded only where you have explicitly consented to the research use described in Section 3(3). By default they are not recorded.
- Information needed for safe operation: to keep use fair and prevent abuse, a per-day request count and amount of processing, recorded against your signed-in account. Rate limiting also uses only that account information; the Service does not read the IP address a request comes from. (Cloudflare, our hosting provider, handles IP addresses as a technical necessity of carrying the traffic — see Section 5.)
The Service uses no advertising trackers. As transient processing needed to compose an answer, your question and the conversation so far are transmitted to the providers listed in Section 5. Transmitting is not the same as storing: what is stored is limited to the items above.
3. Purposes of Use
- Providing and securing the Service
Generating answers, retrieving cited passages, carrying the conversation so far as context, automated screening of input, rate limiting, applying usage limits, and preventing abuse. - Recording usage statistics
To improve the Service and understand how it is used, we store per-request records of the corpus selected, the model, the response mode, processing time, success or failure, and the language setting. These records contain no part of your questions or the answers, no account identifier, and no IP address. Anything we publish is limited to aggregate figures from which no individual can be identified. - Research use of conversation records
Analysing how and for what purposes the Service is used, improving the system on that basis, and publishing the analysis as papers and conference presentations. For this purpose we record your questions, the generated answers, and the sources cited.
This happens only where you have explicitly permitted it in the settings; the default is “do not permit”. Without that permission, no part of your questions or the answers is ever stored, and every feature of the Service remains available to you. You may withdraw consent at any time; on withdrawal recording stops, and you choose then whether to keep or delete the records already stored. Before anything is recorded, email addresses, telephone numbers, and similar items are automatically replaced with placeholders. Records are retained for 730 days from collection. - Billing (once paid plans are offered)
Managing your subscription, processing payment, and keeping the transaction records required by tax law.
Transaction records are kept for the period required by law even if you delete your account. This paragraph applies from the date paid plans begin; the Service currently has no paid plan.
4. The Nature of the Conversation Records
The records described in Section 3(3) contain no account identifier, no email address, and no IP address. They carry only an internal research identifier, and the table mapping that identifier to an account is held separately from the records themselves. Because questions are free text, however, the possibility that an individual could be inferred from what is written cannot be entirely excluded. We therefore do not describe these records as “anonymised information”, and we do not treat them as either anonymously processed information (Art. 43) or pseudonymously processed information (Art. 41) under Japan's Act on the Protection of Personal Information.
We never publish individual conversation records as they stand. What we publish is limited to aggregate figures, the results of statistical analysis, and illustrative examples altered so that no individual can be inferred.
5. External Providers (Processors) and Cross-Border Transfers
The Service relies on the following providers. Each acts as a processor on our instructions; none is permitted to use user information for its own purposes.
| Provider | Role | What is sent | Country of processing |
|---|---|---|---|
| Cloudflare | Hosting, rate limiting, operational logs | Every request, including the originating IP address | United States and others (the location nearest to you) |
| Google (Firebase Authentication / Cloud Firestore) | Sign-in, data storage | Account information and the records described in Section 2 | United States |
| Google (Gemini API, paid tier) | Answer generation, translation of the retrieval query, follow-up suggestions, embedding vectors | Your question, the conversation so far, the retrieved source passages, and the generated answer | United States |
| Pinecone | Vector search and reranking | The text of the translated retrieval query and the text of the candidate passages — not vectors alone | United States |
| OpenAI | Automated content moderation | Your question | United States |
Where a provider offers a setting or contractual term that excludes data from model training, we apply it. Answer generation runs on the paid tier of the API.
All of the above involve handling outside Japan. For the purposes of Art. 28 of Japan's Act on the Protection of Personal Information, we entrust processing only after confirming that the provider maintains a framework equivalent to the standards of that Act, and we verify that this remains so at least once a year. Our agreements with each provider include standard data-protection clauses covering international transfers. Requests for information about the data protection regime of a destination country may be addressed to the contact in Section 1.
6. Retention
| Record | Retention period |
|---|---|
| Usage statistics (Section 3(2)) | Indefinite (they contain no identifier) |
| Research conversation records (Section 3(3)) | 730 days from collection |
| Account information | Until you delete your account (deletable at any time) |
| Per-day request counts | 90 days |
| Rate-limiting counters | 60 seconds |
| Cloudflare operational logs | Cloudflare's default (we cannot change it) |
| Transaction records (once paid plans are offered) | The period required by law (approximately 7 years) |
7. Your Rights
To exercise the rights below, use the contact form in Section 1. We verify your identity and respond within one month as a rule, free of charge.
- Access, rectification, restriction, and erasure (Arts. 33–35 of Japan's Act on the Protection of Personal Information). Account deletion and a download of your own data can also be requested from the settings inside the Service. The download is provided as machine-readable JSON (data portability). Note that we cannot delete your Google account itself; deletion of the Firebase authentication record is performed manually by us, and we confirm to you when it is done.
- Withdrawal of consent. Switching the toggle in the settings back off is all that is required, and no further conversations are recorded. At that point you are asked whether to keep the research conversation records already stored or delete all of them. Records you keep can still be deleted at any time from the account screen. Withdrawal does not affect the lawfulness of processing carried out before it.
- Objection to processing based on legitimate interests. You may object to the processing described in Sections 3(1) and 3(2). Because the usage statistics in Section 3(2) contain no identifier, however, we cannot identify any particular record as yours. In that situation we may be unable to act on the request unless you supply additional information. This limitation is not a design choice made to avoid requests; it follows from keeping identifiers out of the records.
- Complaint to a supervisory authority. Users in Japan may contact the Personal Information Protection Commission; users elsewhere may contact the data protection authority of their country or region. We would nevertheless be grateful for the chance to address the matter first, at the contact in Section 1.
8. Security Measures
- All traffic is encrypted with TLS.
- Authentication is delegated to Google Firebase Authentication; we hold no passwords.
- Usage statistics carry no identifier and their timestamps are rounded to the day. Research conversation records carry only an internal identifier, whose mapping table is kept separately.
- Before a research conversation record is stored, email addresses, telephone numbers, URLs, long digit strings, and similar items are automatically replaced with placeholders.
- Rate limiting operates on the signed-in account, and the Service does not read the IP address a request comes from.
- Database access is confined to a service account holding the minimum necessary privileges.
- Handling in foreign countries: data in the Service is handled principally in the United States. In addition, traffic passing through Cloudflare's network may be processed transiently at the location nearest to you. For an outline of the data protection regimes of those countries, please write to the contact in Section 1.
9. Conversation History in Your Browser
The inquiry history shown on screen is stored inside your browser (localStorage), and that history database itself is never uploaded to a server. On the server side, however, as described in Sections 2 and 3, usage statistics are recorded for every request, and, where you have consented to research use, the text of the conversation is recorded as well. In addition, so that a follow-up question can be answered, the exchanges so far and the passages they drew on are sent to the server and to the providers in Section 5 with each new question. Clearing your browser history does not delete the server-side records; to have those deleted, make a request under Section 7.
10. You Are Interacting with an AI System
Answers in this Service are generated automatically by an AI system; the counterpart is not human. Generated content may contain errors. Please check the cited sources yourself before relying on it.
11. Governing Language
The Japanese text of this Policy is authoritative. The English, Korean, and Chinese versions are translations provided for convenience; in the event of any discrepancy, the Japanese version prevails.
12. Changes to This Policy
Changes to this Policy will be announced within the Service, and renewed consent will be requested for material changes. The effective date and consent version of this Policy is July 28, 2026 (2026-07-28).

