Dials error model overestimation

I’m trying to process some diffraction data, but the error model is refining to poor values, a=10 b=0.01. CC1/2 and CCano are both good, but the sigma dependent statistics are very poor. When disabling the error model the statistics match with the CCs better. But I am concerned with underestimating the sigmas, as I’m trying to do SAD phasing. These crystals do have a 22% tNCS peak, of ~0.04, 0.5, ~0.03 from phaser.

Anyone have an idea on how to better process this data?

Hello @rbauer

What is the value of ISa? a=10 b=0.01 would result in ISa=1/SQRT(a*b)=3.16 if DIALS uses the same definition of a and b that XDS uses (which I am familiar with). That (namely ISa < 5), in my experience, would point to severe errors in indexing and/or space group determination, rendering the data useless (if X-ray or neutron diffraction; the rule does not hold for electron diffraction). If CC1/2 and CCano are good, then I’d ask “how good?”, and what resolution are we talking about? At low resolution, CC1/2 should be close to 1, and CCano - if you want to phase based on the anomalous signal - should be >0.5. How do you want to make use of the anomalous signal, and do you expect metals to be present?
tNCS should not matter for ISa (and the a,b coefficients); it does matter for phasing.
What exactly do you mean with “… the statistics match with the CCs better” - which CCs and statistics?
Best wishes, Kay

Hi Kay,

Thanks for the reply. When using the error model the ISa is 11.7 (from ISa=1/(a*b)), though this varies with dataset. This data is processed in P212121, and when processing with lower symmetry xtriage and other programs indicate that a P2?2?2? space group is more probable.

Both including and not including the error model CC1/2 is good >0.9 until 3.22A. It seemed to me that the I/sI was fairly low, which is improved by excluding the error model. Similarly, <d"/sig> dropped off at 8.5A (and tended to 0 at higher resolution) while CCano appeared significant until~4A. When disabling the error model it seemed to me that there was better respective agreement.

My interpretation was that something was causing the error model to overestimate the sigmas (in the documentation it suggests a=1 and b=0.02-0.03 for good data, here a=9.3 and b=0.009). These are iodine derivatives, and we want to use MR-SAD phasing here, so I want to get an accurate estimation of the error.

I am not too familiar with these programs, so any guidance is appreciated.

Best, Robert

When using the error model:

d_max d_min #obs #uni mult. %comp <I> <I/sI> r_mrg r_meas r_pim r_anom cc1/2 cc_ano
91.32 5.84 126930 8237 15.41 99.98 929.0 5.8 0.298 0.309 0.078 0.144 0.975 0.527
5.84 4.64 134041 7965 16.83 99.96 778.7 3.4 0.429 0.443 0.107 0.172 0.969 0.406
4.64 4.05 114236 7801 14.64 98.93 848.4 2.8 0.457 0.473 0.122 0.213 0.950 0.288
4.05 3.68 128928 7864 16.39 99.92 616.8 2.0 0.539 0.556 0.136 0.208 0.937 -0.036
3.68 3.42 133846 7831 17.09 100.00 456.9 1.4 0.598 0.617 0.149 0.224 0.912 -0.227
3.42 3.22 128559 7781 16.52 99.54 294.1 0.9 0.695 0.718 0.176 0.270 0.907 -0.290
3.22 3.05 120146 7744 15.51 99.08 202.1 0.6 0.815 0.843 0.212 0.339 0.849 -0.258
3.05 2.92 128513 7751 16.58 99.67 149.4 0.5 0.963 0.994 0.243 0.362 0.881 -0.280
2.92 2.81 132253 7776 17.01 99.85 111.6 0.4 1.129 1.164 0.281 0.403 0.841 -0.230
2.81 2.71 132144 7755 17.04 99.82 84.3 0.3 1.380 1.423 0.343 0.449 0.846 -0.178
2.71 2.63 126739 7739 16.38 99.72 67.4 0.2 1.788 1.845 0.451 0.486 0.832 -0.120
2.63 2.55 91516 7047 12.99 91.05 48.1 0.1 2.359 2.456 0.667 0.631 0.790 -0.084
2.55 2.48 69257 6282 11.02 80.80 38.4 0.1 2.807 2.944 0.862 0.763 0.722 -0.033
2.48 2.42 54255 5473 9.91 70.76 30.3 0.1 3.238 3.415 1.055 0.867 0.662 -0.029
2.42 2.37 41741 4677 8.92 60.45 20.6 0.1 4.921 5.217 1.681 0.999 0.626 -0.025
2.37 2.32 31640 3934 8.04 50.72 16.9 0.0 5.858 6.250 2.120 1.116 0.543 -0.020
2.32 2.27 23825 3264 7.30 42.19 10.0 0.0 18.564 19.965 7.170 1.324 0.387 -0.013

Without the error model:

d_max d_min #obs #uni mult. %comp <I> <I/sI> r_mrg r_meas r_pim r_anom cc1/2 cc_ano
91.32 5.84 102907 8237 12.49 99.98 868.9 45.3 0.229 0.240 0.071 0.228 0.972 0.908
5.84 4.64 115954 7965 14.56 99.96 702.0 25.1 0.341 0.355 0.095 0.265 0.970 0.846
4.64 4.05 100312 7801 12.86 98.93 774.2 21.8 0.382 0.399 0.113 0.314 0.960 0.741
4.05 3.68 118926 7864 15.12 99.92 554.1 15.1 0.486 0.504 0.131 0.298 0.955 0.678
3.68 3.42 127901 7831 16.33 100.00 414.5 11.0 0.565 0.584 0.146 0.293 0.943 0.420
3.42 3.22 125666 7781 16.15 99.54 271.7 7.4 0.693 0.717 0.179 0.322 0.933 0.282
3.22 3.05 118454 7744 15.30 99.08 186.8 5.2 0.830 0.859 0.219 0.376 0.897 0.038
3.05 2.92 127691 7751 16.47 99.67 140.2 4.1 1.002 1.034 0.254 0.373 0.896 -0.091
2.92 2.81 131631 7776 16.93 99.85 104.8 3.3 1.186 1.223 0.296 0.407 0.878 -0.093
2.81 2.71 131845 7755 17.00 99.82 79.5 2.5 1.481 1.527 0.367 0.446 0.873 -0.089
2.71 2.63 126591 7739 16.36 99.72 62.8 2.0 2.013 2.078 0.507 0.489 0.855 -0.066
2.63 2.55 91473 7047 12.98 91.05 44.6 1.3 2.788 2.904 0.791 0.644 0.775 -0.053
2.55 2.48 69238 6282 11.02 80.80 35.5 1.0 3.444 3.613 1.061 0.778 0.697 -0.024
2.48 2.42 54252 5473 9.91 70.76 27.8 0.8 3.944 4.161 1.288 0.876 0.668 -0.008
2.42 2.37 41735 4677 8.92 60.45 18.8 0.6 7.047 7.472 2.411 1.029 0.577 -0.005
2.37 2.32 31629 3934 8.04 50.72 15.3 0.5 8.834 9.428 3.205 1.148 0.500 0.007
2.32 2.27 23825 3264 7.30 42.19 8.7 0.3 -105.328 -113.303 -40.774 1.358 0.290 0.002

Hi @rbauer

the tables are very difficult to read - maybe “Blockquote” in edit mode would help. Also, the editor seems to be swallowing the “average intensity” column heading (between percent completeness and average I/sigI).

I would strongly advise against switching off the error model. The reason is that it just tries to tell you something about data, and tries to make sensible estimates of the errors in the intensities. In this case, the a=10 part tries to tell you that the actual differences between symmetry-related reflections are 10 times higher than expected from counting statistics. This is worrisome if the data are from a photon-counting detector (like Pilatus or Eiger), but may be normal for some other detector types e.g. CCDs. Perhaps it is a sign of radiation damage?

Other than that, the meaning of ISa and CC1/2 is somewhat orthogonal. The former measures the systematic error in your data, whereas the latter depends on the sum of systematic and random error. If the random error is low but the systematic error is high, ISa is low (5 .. 10) but CC1/2 can still be fairly high (like 0.95). So I don’t find it useful to predict values of CC1/2 from ISa, or vice versa.

The spacegroup is a hypothesis until successful refinement. “The proof of the pudding is in the eating” so if you can solve and refine the structure with these data in P212121, that proves that the data processing is correct. If you want a second opinion, process the data with XDS.

Best wishes,

Kay

Hi @rbauer , I would be interested to see the error model details printed in the dials.scale.log, that might help to understand what is going on and how well it has fit the data. The section that looks like this:

Error model details:
Type: basic
Parameters: a = 0.55066, b = 0.03564
Error model formula: σ’² = a²(σ² + (bI)²)
estimated I/sigma asymptotic limit: 50.952

Results of error model refinement. Uncorrected and corrected variances
of normalised intensity deviations for given intensity ranges. Variances
are expected to be ~1.0 for reliable errors (sigmas).
±-------------------------±---------±-----------------------±---------------------+
| Intensity range () | n_refl | Uncorrected variance | Corrected variance |
|--------------------------±---------±-----------------------±---------------------|
| 6264.60 - 1944.83 | 450 | 4.643 | 0.928 |
| 1944.83 - 1465.32 | 450 | 3.361 | 1.024 |
| 1465.32 - 1194.56 | 587 | 2.595 | 1.027 |
| 1194.56 - 687.56 | 2243 | 1.854 | 1.041 |
| 687.56 - 395.75 | 3556 | 1.11 | 1.09 |
| 395.75 - 227.78 | 5532 | 0.754 | 1.149 |
| 227.78 - 131.11 | 7687 | 0.512 | 1.096 |
| 131.11 - 75.46 | 9509 | 0.405 | 1.055 |
| 75.46 - 43.43 | 9557 | 0.325 | 0.95 |
| 43.43 - 24.99 | 5446 | 0.302 | 0.927 |
±-------------------------±---------±-----------------------±---------------------+

These are very weird values to refine the parameters to which makes me wonder if there is an issue upstream

Please could you copy and paste the summary table from the end of scaling and the bit where it reports the error model stuff?

hi @rbauer! thanks for your post. I edited your tables to make visualization easier. best of luck processing your data :slight_smile:

@rbauer are you using the parameter anomalous=True for your scaling job? Your data look very anomalous, so you would want this for good scaling. Also if there is high absorption, you might want to try absorption_level=high as this will help get a better scaling correction and therefore reduce the error model parameters.

@JamesBE-DLS I have been using the anomalous flag for scaling. Also, changing the absorption level doesn’t seem to affect much, I think we were fairly conservative during data collection.

Here is the section:

Scale factors determined during minimisation have now been
applied to all reflections for dataset 0.

A round of outlier rejection has been performed,
16321 outliers have been identified.

84 outliers identified from normalised intensity analysis (E² > 100.0)
Performing a round of error model refinement.
b"\nError model details:\n Type: basic\n Parameters: a = 9.37640, b = 0.00911\n Error model formula: \xcf\x83’\xc2\xb2 = a\xc2\xb2(\xcf\x83\xc2\xb2 + (bI)\xc2\xb2)\n estimated I/sigma asymptotic limit: 11.708\n"
Results of error model refinement. Uncorrected and corrected variances
of normalised intensity deviations for given intensity ranges. Variances
are expected to be ~1.0 for reliable errors (sigmas).
±-------------------------±---------±-----------------------±---------------------+
| Intensity range () | n_refl | Uncorrected variance | Corrected variance |
|--------------------------±---------±-----------------------±---------------------|
| 43111.18 - 10872.12 | 250 | 361.339 | 0.955 |
| 10872.12 - 9000.93 | 250 | 267.32 | 1.003 |
| 9000.93 - 4797.63 | 1963 | 211.268 | 1.139 |
| 4797.63 - 2307.66 | 3797 | 137.907 | 1.021 |
| 2307.66 - 1109.99 | 2992 | 91.678 | 0.801 |
| 1109.99 - 533.91 | 1413 | 66.91 | 0.656 |
| 533.91 - 28.57 | 424 | 42.599 | 0.452 |
±-------------------------±---------±-----------------------±---------------------+

The reflection table variances have been adjusted to account for the
uncertainty in the scaling model

Total time taken: 59.3818s

================================================================================

33.33% of model parameters have significant uncertainty
(sigma/abs(parameter) > 0.5)

Summary of dataset partialities
±-----------------±---------+
| Partiality (p) | n_refl |
|------------------±---------|
| all reflections | 1793288 |
| p > 0.99 | 1767621 |
| 0.5 < p < 0.99 | 15084 |
| 0.01 < p < 0.5 | 8406 |
| p < 0.01 | 2177 |
±-----------------±---------+

Reflections below a partiality_cutoff of 0.4 are not considered for any
part of the scaling analysis or for the reporting of merging statistics.
Additionally, if applicable, only reflections with a min_partiality > 0.95
were considered for use when refining the scaling model.

Thanks @rbauer . From the “uncorrected variance” column, the internal deviation does seem high across the intensity range. I have sometimes seen this for multi-sweep anomalous data, although never fully understood the origin. If this is multiple sweeps being processed together, you can try setting error_model.grouping=individual. I found this helped in those cases - as the different sweeps have decent internal consistency, so individual error models will refine to more usual parameters, and then there is still significant deviation is from sweep to sweep (which you then effectively don’t correct for). But this still indicates some underlying systematic error.