Data Generation
Realistic Values
How value profiles make generated data look like real production data.
Realistic Values
Faker Forge doesn't stop at values that merely match a column's type. It writes a value profile for each column that describes what real production data in that column looks like, and generates rows from it.
Why Profiles
A faker method alone gets the type right but not the meaning. A status VARCHAR(20) column filled with random words is valid SQL, but no real application stores "voluptas" as an order status. A profile captures the meaning instead:
status:completed70%,pending20%,cancelled10%sku: codes likePRD-4821-KXprice: 5–500, most values near the low endemail: built from the same row's first and last nametotal:quantity × unit_pricecity,state,postcode,country: all from the same country
How Profiles Are Created
When data is generated for a table for the first time, an AI pass reads the whole schema, infers the business domain, and writes a profile for each column that can have one. This happens once per schema, not once per row. Rows are then generated locally from the profiles, so large row counts cost no extra AI usage.
Any additional context you give when generating (for example "customers are based in Kenya, prices in KES") steers the profiles.
Profile Strategies
| Strategy | Use it for | Example |
|---|---|---|
| Weighted set | Statuses, types, roles, plans | active 70 / inactive 25 / suspended 5 |
| Pool | Realistic names, titles, product names | 30–50 domain-specific strings |
| Pattern | Codes, SKUs, references | ORD-########, [A-Z]{3}-[0-9]{4} |
| Numeric | Prices, quantities, ages, ratings | min, max, decimals, uniform / normal / skewed |
| Date and time | Timestamps, and dates that follow another column | updated_at after created_at |
| Derived | Values built from other columns in the row | {first_name}.{last_name}@gmail.com |
| Computed | Arithmetic over other columns in the row | {quantity} * {unit_price} |
| Correlated group | People and addresses that must agree | name, gender, city, state, postcode, country |
| Faker | When a Faker method is already realistic | jobTitle() |
Any profile on a nullable column can also set a null rate, for example 10% of middle_name values left empty.
Columns That Keep Exact Rules
Some columns never get a profile, because their values have to follow exact rules:
- Primary keys and auto-increment columns
- Foreign keys (values always come from generated parent rows)
- ENUM and SET columns (only the declared values)
- Password columns (a consistent hashed value)
Safety Checks
Every profile is checked against the column before it is used:
- Values longer than the column's length are dropped.
- Numeric ranges are clamped to the column type and to any
CHECKconstraint. - Decimal places never exceed the column's scale.
- Derived and computed profiles may only reference columns of the same table, with no cycles.
A profile that can't be made safe is discarded, and the column falls back to its faker method. If the AI step itself fails, generation carries on with faker methods, so generation never fails because of profiling.
Refining Profiles
Use the Data Assistant on the schema page to change profiles in plain language, for example "make 80% of orders completed", then preview rows before regenerating. MCP clients can do the same with the profile tools (see MCP Tools Reference).
Profile changes apply the next time data is generated.