> ## Documentation Index
> Fetch the complete documentation index at: https://docs.topsort.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Data Requirements

export const LastUpdated = ({date, lang = "en"}) => {
  const translations = {
    en: "Last updated:",
    es: "Última actualización:",
    pt: "Última atualização:",
    fr: "Dernière mise à jour:",
    de: "Zuletzt aktualisiert:"
  };
  const label = translations[lang] || translations.en;
  return <>
<style>{`
.last-updated-component {
display: inline-flex;
align-items: center;
gap: 8px;
padding: 10px 16px;
border-radius: 8px;
margin-top: 12px;
margin-bottom: 16px;
font-size: 14px;
background-color: rgba(0, 0, 0, 0.05);
border: 1px solid rgba(0, 0, 0, 0.12);
color: rgba(0, 0, 0, 0.75);
line-height: 1;
}

        .last-updated-component svg {
          flex-shrink: 0;
          vertical-align: middle;
        }

        .last-updated-component span {
          display: inline-flex !important;
          align-items: center !important;
          line-height: 1 !important;
        }

        [data-theme="dark"] .last-updated-component {
          background-color: #3a3a3a;
          border: 2px solid #888888;
          color: #ffffff;
        }

        [data-theme="dark"] .last-updated-component svg {
          stroke: #ffffff;
        }
      `}</style>
      <div className="last-updated-component">
        <svg width="16" height="16" viewBox="0 0 24 24" fill="none" stroke="currentColor" strokeWidth="2" strokeLinecap="round" strokeLinejoin="round">
          <circle cx="12" cy="12" r="10" />
          <polyline points="12 6 12 12 16 14" />
        </svg>
        <span>
          <strong style={{
    fontWeight: 600
  }}>{label}</strong> 
          <time dateTime={date}>{date}</time>
        </span>
      </div>
    </>;
};

T-Brain learns from your product information and shopper activity. You can upload data from your existing systems or, if you are a Topsort customer, use datasets already synced from Topsort.

To get started, provide your catalog and at least one month of event history, including all purchases and paid product clicks and impressions.

## Dataset types

A dataset is a named collection of data used to train a model. T-Brain supports three dataset types:

| Dataset type | What it contains |
| - | - |
| Categories | Category names and their place in your catalog hierarchy. |
| Products | Product information, including identifiers, descriptions, brands, and category associations. |
| Events | Shopper activity, including impressions, clicks, and purchases. |

You can name datasets to distinguish their purpose or source, such as `products-main-store` or `events-september`.

## Upload data

1. Go to **Data**.
2. Click **Add dataset**.
3. Select **Categories**, **Products**, or **Events**.
4. Give the dataset a name.
5. Upload your **Parquet files** and follow the instructions.

T-Brain displays the required structure for each dataset type and verifies that your files comply with it.

The tables below explain the columns shown in the upload requirements. Fields marked **Required** must be populated. Other fields provide product information, relationships, or context needed for your use case.

## Categories

Category uploads are snapshots. Each upload should contain the full set of categories.

| Column | Type | Requirement | Description |
| - | - | - | - |
| `category_id` | String | Required | Unique category identifier. Products reference this value through `category_ids`. |
| `name` | String | Required | Category display name. |
| `path` | String | Optional | Dot-separated category hierarchy, with the most general category first. |

When a path is provided, its final segment must equal `category_id`. For example, a category with ID `air-fryers` could have the path `home.kitchen.air-fryers`.

Path segments may contain letters, numbers, underscores, and hyphens. The path can be null for a root category or a catalog without a hierarchy.

## Products

Product uploads are snapshots. Each upload should contain the full set of products.

| Column | Type | Requirement | Description |
| - | - | - | - |
| `product_id` | String | Required | Unique product identifier. Must exactly match the corresponding `product_id` in event records. |
| `name` | String | Optional | Product display name, used to help the model understand the product. |
| `description` | String | Optional | Product description. HTML is accepted and stripped during processing. |
| `image_url` | String | Optional | Product image URL, used when generating image embeddings. |
| `brand_id` | String | Optional | Brand identifier. |
| `brand_name` | String | Optional | Brand display name. |
| `category_ids` | List of strings | Optional | Category identifiers matching records in the categories dataset. |
| `price` | Decimal | Optional | Price in your marketplace’s currency units. Zero is a valid price; null means unknown. |
| `active` | Boolean | Optional | Whether the product is available for selection by the model. |
| `parent_product_id` | String | Optional | Parent product identifier for a variant. |

Although only `product_id` is structurally required, include descriptive information where available so the model has useful information about your products.

## Events

Provide at least one month of event history. Include all purchases, not only purchases attributed to advertising, plus paid product clicks and impressions.

Events must include a consistent shopper identifier so interactions and purchases can be connected. In the dataset, this identifier is stored in `user_id` and must remain stable across sessions and days.

Event uploads are additive, allowing you to add activity from additional periods.

| Column | Type | Requirement | Description |
| - | - | - | - |
| `event_id` | String | Required | Unique event-row identifier. Each purchase line has its own identifier. |
| `event_type` | String | Required | Type of activity. See supported values below. |
| `ts` | Datetime | Required | When the activity occurred, rather than when the file was uploaded. |
| `user_id` | String | Needed to connect shopper activity | Consistent shopper identifier across sessions and days. |
| `product_id` | String | Where applicable | Product involved in the event. Can be null for page-level and request events. |
| `request_id` | String | Where applicable | Request or auction identifier associated with the event. |
| `source` | String | Where applicable | `sponsored` or `organic`. |
| `placement` | String | Optional | Where the interaction occurred on the page. |
| `search_term` | String | Optional | Shopper’s search query, decoded and trimmed. |
| `page_type` | String | Optional | `home`, `category`, `search`, `pdp`, or `other`. `pdp` means product detail page. |
| `device` | String | Optional | `mobile` or `desktop`. |
| `category` | String | Where applicable | Category identifier associated with the request or page. |
| `order_id` | String | For purchase rows | Groups the product lines belonging to the same order. |
| `quantity` | Integer | For purchase rows | Number of units purchased on that line. |
| `unit_price` | Decimal | For purchase rows | Paid price per unit, in your marketplace’s currency units. |
| `product_ids` | List of strings | For request rows | Products shown or considered in the request. |

### Supported event types

| Value | Meaning |
| - | - |
| `impression` | A shopper viewed a product placement. |
| `click` | A shopper clicked a product. |
| `add_to_cart` | A shopper added a product to their cart. |
| `page_view` | A shopper viewed a page. |
| `purchase` | A shopper purchased a product. |

### Keep identifiers consistent

Use exact, consistent identifiers across datasets:

* An event’s `product_id` must match the product dataset’s `product_id`.
* Values in a product’s `category_ids` must match category records’ `category_id`.
* Use the same `user_id` for the same shopper across interactions and purchases.
* Give each purchase line a unique `event_id`, and use `order_id` to group lines from the same order.

## Refresh an uploaded dataset

In **Data**, click **Add Files** next to the dataset you want to update.

For products and categories, provide a full snapshot. For events, add the new event records.

Once the dataset is updated, retrain any model that should learn from the refreshed information. Adding files does not automatically refresh an already trained model.

## Use synced Topsort datasets

If you are a Topsort customer, three datasets will already be available:

* `Topsort-categories`
* `Topsort-products`
* `Topsort-events`

These datasets update in real time. You can select them when creating a task instead of uploading the same information yourself.

Real-time dataset updates do not automatically retrain deployed models.

## Train with your datasets

Go to **Tasks**, choose what you want the model to do, and select the relevant datasets. For example, a product-ranking task uses products and events.

Start training, follow progress in **Training**, and review model quality in **Evals** before deployment.

## Data privacy and access

Your uploaded data and fine-tuned models are isolated per retailer. Retention and deletion policies can be configured based on your requirements.

[Talk to a Topsort sales representative](https://www.topsort.com/book-a-demo) to arrange access. API documentation is available on request, including authentication, request and response formats, latency, and rate limits.

***

<LastUpdated date="2026-10-08" />


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.