Legal FAQ

These answers describe self-serve use of Jsonify. Enterprise customers have a signed agreement that sets their own terms, roles and scope; where an answer says "enterprise", that agreement governs.

Is Jsonify a web scraping tool?

Jsonify collects data from public websites and mobile apps by automated means. That is web scraping in the ordinary sense of the term. What you buy is a pipeline built from your brief, checks on every row, a dataset published to your tools on schedule, and repair when a source changes. The legal questions are the same as for any automated collection: which sources, what access, what data, and what use. You confirm the sources and fields before anything is collected. Jsonify does not access content behind a login or paywall. The answers below cover the rest.

Who decides what is collected?

You do. Jason proposes sources, fields and a schedule; nothing runs until you confirm them. Jason may later suggest new sources or columns; those are applied only if you accept them. Your confirmed configuration is the instruction Jsonify works to.

What types of data does Jsonify collect?

Publicly available commercial information from websites and mobile apps: product listings, prices, availability, menus, catalogs, and similar data. Pipelines can also read files you upload as sources. Collection targets products and services, not people. Self-serve pipelines read only public content; access to anything requiring authentication exists only under an enterprise agreement, using credentials the customer legitimately holds.

How does Jsonify handle robots.txt and access controls?

For websites, robots.txt is read when a source is added to a pipeline, and no collection is built for disallowed paths; self-serve users cannot override this. robots.txt is a convention for websites and does not apply to mobile apps, where Jsonify reads only what the app shows without signing in. Jsonify does not access content behind logins, paywalls or subscriptions, and does not use your credentials on self-serve plans. Pipelines limit request rates and cache fetched content so sources are requested fewer times. Jsonify chooses the technical method of collection, including the networks and browsers requests are sent through.

Does Jsonify agree to a website's terms of use on my behalf?

No. Self-serve collection does not tick boxes, register accounts or submit forms, so Jsonify does not accept any source's terms for you. Whether a source's terms apply to you and what they permit is for you to assess. Products that complete forms (Benchmark) are enterprise-only and governed by that agreement.

Who is responsible for legal compliance when using Jsonify?

Jason proposes sources and schemas, but nothing runs until you confirm it, and you are responsible for what you confirm: whether you are authorized to access and use the data, and whether your use complies with applicable laws and any source terms that bind you. Jsonify is responsible for operating the Services as described in the Terms — respecting robots.txt on websites, staying out of gated content, and processing personal data only as your processor or as set out in a signed agreement.

What are the data protection roles in self-serve and enterprise?

On self-serve plans, you are the data controller for the pipelines you configure and Jsonify is your data processor under the Data Processing Addendum, which is incorporated into the Terms. Under enterprise agreements, the roles are as set out in the signed agreement, which takes precedence over the public Terms; some enterprise agreements are structured with the parties as independent controllers.

Is customer data used to train models? Which AI providers are involved?

No. Jsonify does not use customer data to train models, and does not copy rows delivered to a customer into datasets offered to others. Jsonify may collect the same public data independently, for itself or for other customers; that collection is its own, not customer data. Pipelines are built and repaired using third-party large language models acting as subprocessors, under terms that do not permit them to train on the data. The current providers are listed on the subprocessors page. Jsonify may publish de-identified operational metrics about the platform, such as repair times, that do not reveal customer data.

How does Jsonify comply with the EU AI Act?

Jason is an AI system, and Jsonify discloses that wherever you interact with it — in the product and in communications Jason sends — in line with the transparency obligations of Article 50 of the EU AI Act, which have applied since August 2026. Model-derived fields in datasets can be wrong; the Terms say so, and every row links to its source page so you can verify it.

Is personal data anonymised?

Pseudonymisation and filtering of personal-data fields can be configured into a pipeline under an enterprise agreement; where configured, it runs on every row. It is not enabled on self-serve pipelines. Note that techniques such as hashing or masking usernames are pseudonymisation rather than anonymisation in the GDPR sense; either way, personal data in your pipelines is handled under the Data Processing Addendum.

Where is data hosted?

Datasets and run artifacts are stored on infrastructure in the European Union. Processing also happens in the United States through the fetch and model providers listed on the subprocessors page. International transfers are covered by the safeguards described in the Privacy Policy.

What happens if a website or app owner objects to collection?

Jsonify may stop collecting from any source at any time, for all customers, including when the source's operator asks. Affected customers are notified where practicable and are not billed for rows not collected. Site and app owners: see For website and app owners.

What is a row, and how does billing work?

A row is counted only when it has passed its checks and reached your dataset; every run counts. Each month up to 100 rows are free; above that, usage is billed at month end in arrears with a US$49 monthly minimum, at the rates on the pricing page. Builds, previews, retries, blocked pages and repairs are never charged. Billing starts only after you explicitly enable it and add a payment method. The full terms are in the Self-Serve Terms.

Can I use the data for decisions about people?

No. The Acceptable Use Policy prohibits using datasets to make or support decisions about an individual's credit, employment, housing, insurance or similar, and prohibits profiling, locating or monitoring individuals.

What do we own, and what does Jsonify own?

Your schemas, source lists and collected rows are yours, subject to third-party rights. Jsonify owns the platform, its reusable extractors for public sources and the technical knowledge of how public pages are structured, and reuses those across customers. Raw fetched content may be cached and shared across pipelines to reduce load on sources; your rows, schemas and source lists are never shown to other customers. Enterprise code export is governed by the licence in your agreement. Read the ownership terms →

Connect your data assistant

Build datasets and work with your data in Claude, ChatGPT, Copilot or another assistant.

Connect in Claude

  1. Open the Jsonify connector: Open on Claude.ai ↗
  2. Select Connect and sign in to Jsonify.
  3. In a chat, enable Jsonify under + → Connectors and describe the data you need.
Advanced

Team or Enterprise account? Your workspace owner enables Jsonify once for everyone from Admin settings → Connectors. Jsonify is a community connector in the directory — some admins allow those by default, some review them first.

No directory? Add it by hand: Settings → Connectors → + Add custom connector, name it Jsonify, paste the server URL below, then Connect.

Server URLhttps://factory.jsonify.com/mcp

Official Claude custom-connector guide ↗

Then say: “build me a dataset of competitor product prices and availability, refreshed daily”. Full instructions per client on /connect.